llmfit: choosing local models based on your real hardware
AlexsJones/llmfit · 37,154★ · 2,368 forks
Everything worth knowing about AlexsJones/llmfit: a terminal tool and an interactive interface that detects a machine and ranks local models by fit, estimated speed, quality, and context.
What llmfit is
llmfit helps decide which local language model can run well on a specific machine. It detects RAM, CPU, GPU, VRAM, and installed providers; it then scores the catalog by memory fit, speed, quality, and context. It offers an interactive terminal interface by default and a command-line mode for automation.

The repository documents compatibility with Ollama, llama.cpp, MLX, Docker Model Runner, and LM Studio, plus multi-GPU setups, MoE architectures, and dynamic quantization selection. It doesn’t necessarily run every candidate to recommend it: its estimates start from specifications and expose their assumptions via llmfit info; the benchmark suite lets you replace estimates with your own measurements.

The origin: from justifying a laptop to cataloging fit
Alex Jones created the repository on February 15, 2026. Their profile describes them as a principal engineer at AWS, with their work described as infrastructure for the age of agents.

The launch recovered on Hacker News, submitted by the account axjns on February 17, 2026, gives the most direct context: its author wrote that they built it to justify buying a more powerful laptop. The submission, thread 47045434, had 1 point and 0 comments in the index consulted; it establishes the initial presentation, not an independent reception.
Philosophy and principles
The approach favors verifiable decisions over a generic model list: the tool shows the fit and the basis for its estimates, and points the user toward verifying on their own machine. The community workflow continues that idea: a local measurement is saved first and only optionally shared via a pull request; accepted measurements feed into future releases for identical machines.
There’s also a practical separation between recommendation and experimentation. Estimating doesn’t require downloading or starting every model; to confirm performance, bench measures tokens per second and time-to-first-token against a provider that’s already running.
How it works
- On startup, llmfit inspects local hardware and runtimes; it computes memory fit, estimated speed, quality, and context for each catalog entry.
- The interface shows sorted results and lets you search, download via a provider, and check model details. The classic mode returns tables or JSON for scripts and agents.

planturns a requested configuration — model, context, quantization, and speed target — into hardware requirements and possible execution paths.benchconnects to a running provider, keeps runs locally, and can open a pull request with the data, after confirmation and GitHub authentication.

The codebase is a Rust workspace with llmfit-core for detection and analysis, llmfit-tui for the interface, and llmfit-desktop for the desktop app; it also contains Python bindings and a web interface.
Official and semi-official status
No evidence was found of acceptance into an official marketplace by an AI provider, nor of a manufacturer certification. There are concrete distribution channels documented by the project: a Homebrew formula, MacPorts, Scoop, a uv or pip package, the ghcr.io/alexsjones/llmfit image, and a crates.io package.
This implies availability through those channels, not a compatibility guarantee or formal standard designation. The 31,374 stars visible via the API and the 2,344 recent downloads recorded by crates.io suggest considerable attention, but don’t equate to active installs or unique users.
The ecosystem
Sibling projects and extensions
The README calls the following sibling projects:
sympozium-ai/sympozium: agent management on Kubernetes, per the relationship the project declares.AlexsJones/llmserve: an interface for choosing a model, an engine, and serving a local model; the GitHub search returned 338 stars.AlexsJones/llama-panel: a macOS app for managing localllama-serverinstances; the search returned 55 stars.

Integration with OpenClaw has its own documentation in the repository tree, and the command-line interface exposes an HTTP server with fit data and model selection for cluster schedulers or aggregators.
Forks, translations, and derivatives
The README itself offers translations into Chinese and Japanese (README.zh.md and README.ja.md). These are translations included in the main repository, not separate ports.
A query of the most prominent forks returned copies like smirk-dev/llmfit and abdennour/llmfit, each with 6 stars and the same description as the original. Without documentation proving substantive changes, they’re classified as forks rather than independent extensions.
The README identifies Pavelevich/llm-checker as an alternative: it’s a Node.js tool with Ollama integration that downloads and tests models directly, while llmfit estimates from specifications and supports MoE architectures.
Repo numbers
Measured: August 13, 2026; GitHub API and crates.io.
| Metric | Value |
|---|---|
| Stars | 31,374 |
| Forks | 1,928 |
| Real subscribers | 91 |
| Commits visible on page | 1,075 |
| Open issues per API | 57 |
| Primary language | Rust |
| License | MIT |
| Created | February 15, 2026 |
| Latest release found | v1.1.9, August 9, 2026 |
| Total downloads on crates.io | 16,677 |
| Recent downloads on crates.io | 2,344 |
The top contributors returned by the API were AlexsJones (631 contributions), dependabot[bot] (74), three-foxes-in-a-trenchcoat (44), and github-actions[bot] (29). The API returns open_issues_count, a field that may include open pull requests; it therefore doesn’t necessarily equal an issues-only count. Also, watchers_count mirrors stars in GitHub’s general API response: subscribers_count is reported here as the real subscriber count.
How to contribute
The guide requires stable Rust with the 2024 edition and MSRV 1.85 or later, Python 3 for model-database scripts, and Git. The contribution flow is: fork the repository, branch from main, make changes, run cargo fmt, make clippy, and make test, and open a clear pull request focused on one fix or feature.
To add a model, the guide says to add its Hugging Face identifier to TARGET_MODELS in scripts/scrape_hf_models.py, run make update-models, validate with ./target/release/llmfit list, update MODELS.md if applicable, and open the pull request.

How the community received it
The recoverable direct evidence on Hacker News is limited. The author’s announcement in thread 47045434 got 1 point and no comments, so it contains no verifiable third-party praise or criticism.
The HN index also surfaced a later mention by clmnt: they introduced hf-agents and stated they use llmfit to profile hardware and select a model and quantization before launching Pi Agent. The available recovery didn’t provide the figures or the parent thread for that mention; it’s a statement from another project’s author, not an independent evaluation.
As an operational signal, open issues include a request to share benchmarks (issue 867) and a query about sharing results (issue 862), while the documentation explains that results are saved locally before being published. This shows interest in the measurement workflow, not a quality judgment from its authors.
Queries of Reddit, X, Product Hunt, video, third-party blogs, and podcasts produced no recoverable and verifiable evidence during this run. No opinions, launches, or figures are therefore attributed to those platforms.
llmfit versus other approaches
| Approach | Verifiable overlap | Verifiable difference |
|---|---|---|
Pavelevich/llm-checker | Both help work with local models and relate to Ollama. | llmfit’s README describes llm-checker as downloading and testing models directly; llmfit can estimate from specifications and accounts for MoE. |
| Ollama, llama.cpp, MLX, Docker Model Runner, and LM Studio | These are runtimes llmfit detects or queries. | They aren’t one-to-one substitutes: llmfit uses them as providers to discover, serve, or measure models. |
AlexsJones/llmserve | Both target local models and have an interactive interface. | llmserve presents itself as model and engine selection to serve it; llmfit focuses on fit, recommendations, and hardware planning. |
Quick-start guide
Installation and first run
On macOS or Linux, the recommended binary install is brew install AlexsJones/llmfit/llmfit; on Windows, scoop install llmfit. Also documented are port install llmfit, uv tool install -U llmfit, uvx llmfit, the container image, and building with cargo build --release.

After installing, running llmfit opens the interactive interface and shows detected specs and ranked models. For diagnostics before reporting an issue, use llmfit doctor; for a non-interactive result, llmfit fit or llmfit recommend --json.
Common workflows
- Choosing models for coding:
llmfit recommend --json --use-case coding --limit 3returns filtered recommendations for automation. - Checking a candidate:
llmfit info "Mistral-7B"shows the fit analysis, the estimate bases, and verification commands. - Planning a purchase or a server:
llmfit plan "Qwen/Qwen3-4B-MLX-4bit" --context 8192 --target-tps 25 --jsoncomputes requirements and execution alternatives. - Measuring and sharing: with an active provider,
llmfit bench --all --share --dry-runpreviews the data;llmfit bench --all --sharemeasures and offers to open a pull request.
Essential configuration
--memory,--ram, and--cpu-cores: override VRAM, RAM, or core detection for virtual machines or simulations.--max-context: limits the context used in the memory estimate; without it, it may useOLLAMA_CONTEXT_LENGTH.--force-runtime: forcesmlx,llamacpp, orvllmin the analysis when automatic selection isn’t appropriate.LLAMA_SERVER_HOSTorLLAMA_SERVER_PORT: change llama.cpp detection forbench.LLMFIT_BENCH_STORE: changes the location of pending measurements on Linux.
Common pitfalls and fixes
- If automatic detection fails on a virtual machine or with a broken
nvidia-smi, use the hardware overrides, for examplellmfit --memory=24G --ram=64G fit. - A benchmark needs a running provider; for llama.cpp, the guide shows
llama-server -m ~/.cache/llmfit/models/Qwen3.5-0.8B-Q8_0.gguf --port 8080 -ngl 99and thenrin the interface to refresh installed models. --shareauthenticates before measuring and fails early if the credential is missing or expired. Use the device flow or setGITHUB_TOKENorGH_TOKEN;--dry-runavoids touching the network.- The default GGUF download directory is
~/.cache/llmfit/models; it can be changed from the interface’s download-manager settings.
Integrations and migration
llmfit can output JSON for scripts and agents, or run llmfit serve --host 0.0.0.0 --port 8787 to expose /health, /api/v1/system, and model selection via HTTP to cluster schedulers. For workflows already built on Ollama, llama.cpp, MLX, Docker Model Runner, or LM Studio, the documented pattern is to keep that provider and let llmfit detect it for inventory or benchmarks.
Use cases
- Anyone wanting to run models locally on a laptop or workstation can compare memory fit and an estimated speed before downloading large models.
- Teams managing multiple GPUs, Apple Silicon, or uneven hardware can override detection and simulate RAM, VRAM, cores, and context to plan a reproducible configuration.
- People responsible for automation or cluster planning can consume JSON or the HTTP API to choose the models runnable on each node.
- People maintaining local providers can measure Ollama, vLLM, MLX, or llama.cpp, keep their data locally first, and optionally contribute it so identical machines benefit from a later release.
Resources
- Repository: https://github.com/AlexsJones/llmfit
- Documentation and installation: https://github.com/AlexsJones/llmfit#install
- Interface, benchmarks, and usage guide: https://github.com/AlexsJones/llmfit/tree/main/docs
- Rust package: https://crates.io/crates/llmfit
- Sibling project for serving models: https://github.com/AlexsJones/llmserve
- Hacker News discussion: https://news.ycombinator.com/item?id=47045434
- Community: https://github.com/AlexsJones/llmfit/discussions
Note: this article combines documentation, the contribution guide, the GitHub API, crates.io, and the Hacker News index retrieved on August 13, 2026. Figures change over time.
Comments