August 23, 2026 · By YasKad
AlexsJones/llmfit

llmfit: choosing local models based on your real hardware

AlexsJones/llmfit · 37,154★ · 2,368 forks

Everything worth knowing about AlexsJones/llmfit: a terminal tool and an interactive interface that detects a machine and ranks local models by fit, estimated speed, quality, and context.


What llmfit is

llmfit helps decide which local language model can run well on a specific machine. It detects RAM, CPU, GPU, VRAM, and installed providers; it then scores the catalog by memory fit, speed, quality, and context. It offers an interactive terminal interface by default and a command-line mode for automation.

Ultra-detailed dark cyberpunk illustration of a glowing, neon-outlined computer motherboard. RAM modules and a massive GPU are lit up in vibrant cyan and magenta, with holographic data streams scanning across the hardware components, representing the hardware detection and system inspection phase.

The repository documents compatibility with Ollama, llama.cpp, MLX, Docker Model Runner, and LM Studio, plus multi-GPU setups, MoE architectures, and dynamic quantization selection. It doesn’t necessarily run every candidate to recommend it: its estimates start from specifications and expose their assumptions via llmfit info; the benchmark suite lets you replace estimates with your own measurements.

Ultra-detailed dark-mode cyberpunk visualization of a futuristic data marketplace. Abstract holographic containers labeled with various AI model names are being sorted and scanned by robotic arms emitting neon laser beams, with the models being ranked on a glowing neon leaderboard.

The origin: from justifying a laptop to cataloging fit

Alex Jones created the repository on February 15, 2026. Their profile describes them as a principal engineer at AWS, with their work described as infrastructure for the age of agents.

Ultra-detailed cyberpunk-inspired 3D render of a sleek laptop emitting a massive glowing holographic neural network. The hologram forms a perfect sphere of interconnected neon nodes, floating above the laptop screen, bathed in deep neon blue and purple light, representing the tool's origin: justifying the purchase of a powerful laptop to run local AI models.

The launch recovered on Hacker News, submitted by the account axjns on February 17, 2026, gives the most direct context: its author wrote that they built it to justify buying a more powerful laptop. The submission, thread 47045434, had 1 point and 0 comments in the index consulted; it establishes the initial presentation, not an independent reception.

Philosophy and principles

The approach favors verifiable decisions over a generic model list: the tool shows the fit and the basis for its estimates, and points the user toward verifying on their own machine. The community workflow continues that idea: a local measurement is saved first and only optionally shared via a pull request; accepted measurements feed into future releases for identical machines.

There’s also a practical separation between recommendation and experimentation. Estimating doesn’t require downloading or starting every model; to confirm performance, bench measures tokens per second and time-to-first-token against a provider that’s already running.

How it works

  1. On startup, llmfit inspects local hardware and runtimes; it computes memory fit, estimated speed, quality, and context for each catalog entry.
  2. The interface shows sorted results and lets you search, download via a provider, and check model details. The classic mode returns tables or JSON for scripts and agents.

Ultra-detailed dark-mode cyberpunk conceptual UI illustration: a futuristic command-line terminal window floating in a dark void, displaying neon-green text and structured JSON data tables. The text glows softly, casting a neon reflection onto an abstract metallic surface below, representing CLI mode and JSON output for automation and agents.

  1. plan turns a requested configuration — model, context, quantization, and speed target — into hardware requirements and possible execution paths.
  2. bench connects to a running provider, keeps runs locally, and can open a pull request with the data, after confirmation and GitHub authentication.

Visually rich dark-mode cyberpunk scene depicting a futuristic benchmarking lab. A glowing CPU and GPU enclosed in transparent glass are connected by fiber-optic cables pulsing with neon orange and cyan light. Digital speedometers and tokens-per-second metrics float in the air as holographic displays, representing the bench command measuring local model performance.

The codebase is a Rust workspace with llmfit-core for detection and analysis, llmfit-tui for the interface, and llmfit-desktop for the desktop app; it also contains Python bindings and a web interface.

Official and semi-official status

No evidence was found of acceptance into an official marketplace by an AI provider, nor of a manufacturer certification. There are concrete distribution channels documented by the project: a Homebrew formula, MacPorts, Scoop, a uv or pip package, the ghcr.io/alexsjones/llmfit image, and a crates.io package.

This implies availability through those channels, not a compatibility guarantee or formal standard designation. The 31,374 stars visible via the API and the 2,344 recent downloads recorded by crates.io suggest considerable attention, but don’t equate to active installs or unique users.

The ecosystem

Sibling projects and extensions

The README calls the following sibling projects:

  • sympozium-ai/sympozium: agent management on Kubernetes, per the relationship the project declares.
  • AlexsJones/llmserve: an interface for choosing a model, an engine, and serving a local model; the GitHub search returned 338 stars.
  • AlexsJones/llama-panel: a macOS app for managing local llama-server instances; the search returned 55 stars.

Ultra-detailed cyberpunk digital landscape showing a complex network of interconnected neon nodes forming an abstract neural network. The network is organized into distinct glowing clusters, representing different local runtimes such as Ollama, llama.cpp, MLX, and Docker. The dark background is filled with faintly scrolling code.

Integration with OpenClaw has its own documentation in the repository tree, and the command-line interface exposes an HTTP server with fit data and model selection for cluster schedulers or aggregators.

Forks, translations, and derivatives

The README itself offers translations into Chinese and Japanese (README.zh.md and README.ja.md). These are translations included in the main repository, not separate ports.

A query of the most prominent forks returned copies like smirk-dev/llmfit and abdennour/llmfit, each with 6 stars and the same description as the original. Without documentation proving substantive changes, they’re classified as forks rather than independent extensions.

The README identifies Pavelevich/llm-checker as an alternative: it’s a Node.js tool with Ollama integration that downloads and tests models directly, while llmfit estimates from specifications and supports MoE architectures.

Repo numbers

Measured: August 13, 2026; GitHub API and crates.io.

MetricValue
Stars31,374
Forks1,928
Real subscribers91
Commits visible on page1,075
Open issues per API57
Primary languageRust
LicenseMIT
CreatedFebruary 15, 2026
Latest release foundv1.1.9, August 9, 2026
Total downloads on crates.io16,677
Recent downloads on crates.io2,344

The top contributors returned by the API were AlexsJones (631 contributions), dependabot[bot] (74), three-foxes-in-a-trenchcoat (44), and github-actions[bot] (29). The API returns open_issues_count, a field that may include open pull requests; it therefore doesn’t necessarily equal an issues-only count. Also, watchers_count mirrors stars in GitHub’s general API response: subscribers_count is reported here as the real subscriber count.

How to contribute

The guide requires stable Rust with the 2024 edition and MSRV 1.85 or later, Python 3 for model-database scripts, and Git. The contribution flow is: fork the repository, branch from main, make changes, run cargo fmt, make clippy, and make test, and open a clear pull request focused on one fix or feature.

To add a model, the guide says to add its Hugging Face identifier to TARGET_MODELS in scripts/scrape_hf_models.py, run make update-models, validate with ./target/release/llmfit list, update MODELS.md if applicable, and open the pull request.

Striking dark-mode cyberpunk illustration of a futuristic developer workspace. Multiple holographic terminal windows float in the air, displaying Rust programming code and GitHub pull request diffs. Neon green and purple accents highlight the syntax, while a glowing Rust gear logo is subtly integrated into the background.

How the community received it

The recoverable direct evidence on Hacker News is limited. The author’s announcement in thread 47045434 got 1 point and no comments, so it contains no verifiable third-party praise or criticism.

The HN index also surfaced a later mention by clmnt: they introduced hf-agents and stated they use llmfit to profile hardware and select a model and quantization before launching Pi Agent. The available recovery didn’t provide the figures or the parent thread for that mention; it’s a statement from another project’s author, not an independent evaluation.

As an operational signal, open issues include a request to share benchmarks (issue 867) and a query about sharing results (issue 862), while the documentation explains that results are saved locally before being published. This shows interest in the measurement workflow, not a quality judgment from its authors.

Queries of Reddit, X, Product Hunt, video, third-party blogs, and podcasts produced no recoverable and verifiable evidence during this run. No opinions, launches, or figures are therefore attributed to those platforms.

llmfit versus other approaches

ApproachVerifiable overlapVerifiable difference
Pavelevich/llm-checkerBoth help work with local models and relate to Ollama.llmfit’s README describes llm-checker as downloading and testing models directly; llmfit can estimate from specifications and accounts for MoE.
Ollama, llama.cpp, MLX, Docker Model Runner, and LM StudioThese are runtimes llmfit detects or queries.They aren’t one-to-one substitutes: llmfit uses them as providers to discover, serve, or measure models.
AlexsJones/llmserveBoth target local models and have an interactive interface.llmserve presents itself as model and engine selection to serve it; llmfit focuses on fit, recommendations, and hardware planning.

Quick-start guide

Installation and first run

On macOS or Linux, the recommended binary install is brew install AlexsJones/llmfit/llmfit; on Windows, scoop install llmfit. Also documented are port install llmfit, uv tool install -U llmfit, uvx llmfit, the container image, and building with cargo build --release.

Highly detailed dark-mode cyberpunk 3D render of a futuristic toolkit. Floating holographic icons represent different package managers (Homebrew, Scoop, Docker, pip) arranged in a circular pattern around a central glowing neon core resembling an AI chip. Bright data streams connect the icons to the core.

After installing, running llmfit opens the interactive interface and shows detected specs and ranked models. For diagnostics before reporting an issue, use llmfit doctor; for a non-interactive result, llmfit fit or llmfit recommend --json.

Common workflows

  • Choosing models for coding: llmfit recommend --json --use-case coding --limit 3 returns filtered recommendations for automation.
  • Checking a candidate: llmfit info "Mistral-7B" shows the fit analysis, the estimate bases, and verification commands.
  • Planning a purchase or a server: llmfit plan "Qwen/Qwen3-4B-MLX-4bit" --context 8192 --target-tps 25 --json computes requirements and execution alternatives.
  • Measuring and sharing: with an active provider, llmfit bench --all --share --dry-run previews the data; llmfit bench --all --share measures and offers to open a pull request.

Essential configuration

  • --memory, --ram, and --cpu-cores: override VRAM, RAM, or core detection for virtual machines or simulations.
  • --max-context: limits the context used in the memory estimate; without it, it may use OLLAMA_CONTEXT_LENGTH.
  • --force-runtime: forces mlx, llamacpp, or vllm in the analysis when automatic selection isn’t appropriate.
  • LLAMA_SERVER_HOST or LLAMA_SERVER_PORT: change llama.cpp detection for bench.
  • LLMFIT_BENCH_STORE: changes the location of pending measurements on Linux.

Common pitfalls and fixes

  • If automatic detection fails on a virtual machine or with a broken nvidia-smi, use the hardware overrides, for example llmfit --memory=24G --ram=64G fit.
  • A benchmark needs a running provider; for llama.cpp, the guide shows llama-server -m ~/.cache/llmfit/models/Qwen3.5-0.8B-Q8_0.gguf --port 8080 -ngl 99 and then r in the interface to refresh installed models.
  • --share authenticates before measuring and fails early if the credential is missing or expired. Use the device flow or set GITHUB_TOKEN or GH_TOKEN; --dry-run avoids touching the network.
  • The default GGUF download directory is ~/.cache/llmfit/models; it can be changed from the interface’s download-manager settings.

Integrations and migration

llmfit can output JSON for scripts and agents, or run llmfit serve --host 0.0.0.0 --port 8787 to expose /health, /api/v1/system, and model selection via HTTP to cluster schedulers. For workflows already built on Ollama, llama.cpp, MLX, Docker Model Runner, or LM Studio, the documented pattern is to keep that provider and let llmfit detect it for inventory or benchmarks.

Use cases

  • Anyone wanting to run models locally on a laptop or workstation can compare memory fit and an estimated speed before downloading large models.
  • Teams managing multiple GPUs, Apple Silicon, or uneven hardware can override detection and simulate RAM, VRAM, cores, and context to plan a reproducible configuration.
  • People responsible for automation or cluster planning can consume JSON or the HTTP API to choose the models runnable on each node.
  • People maintaining local providers can measure Ollama, vLLM, MLX, or llama.cpp, keep their data locally first, and optionally contribute it so identical machines benefit from a later release.

Resources


Note: this article combines documentation, the contribution guide, the GitHub API, crates.io, and the Hacker News index retrieved on August 13, 2026. Figures change over time.

Comments