~/wiki

llmfit

Mis à jour le 2026-04-14Confiance : high
llmfitmodel-selectionhardware-compatibilityopen-source-toolautomationquantizationinference-optimization

Open-source command-line tool created by eric-vyacheslav that automatically matches large language models to hardware capabilities, solving the common problem of downloading models that won't run on available systems.

Core Functionality

Hardware Profiling

llmfit performs comprehensive system analysis:

CPU Analysis

  • Processor model and architecture detection
  • Core count and thread capabilities
  • Instruction set support (AVX, etc.)

Memory Assessment

  • Total system RAM
  • Available memory (accounting for OS and applications)
  • Memory bandwidth characteristics

GPU Detection

  • Graphics card model and capabilities
  • VRAM size and availability
  • Compute capability and driver support

Storage Analysis

  • Available disk space for model storage
  • Storage type (SSD vs HDD) for loading speed

Multi-Dimensional Scoring

llmfit evaluates each model across four key dimensions:

1. Quality Score

  • Based on parameter count (higher generally means better capability)
  • Quantization impact on model performance
  • Architecture-specific quality factors

2. Speed Score

  • Estimated tokens per second for user's specific hardware
  • Backend optimization considerations
  • CPU vs GPU execution predictions

3. Fit Score

  • Model memory requirements vs available resources
  • Safety margins for stable operation
  • Activation and KV cache overhead

4. Context Window Score

  • Long conversation support within memory constraints
  • Context length vs memory usage trade-offs

Model Compatibility Labels

Perfect Fit

  • Model runs optimally within hardware constraints
  • Excellent performance expected
  • Strongly recommended

Good Fit

  • Model runs well with acceptable trade-offs
  • Good performance expected
  • Viable deployment option

Marginal Fit

  • Model operates at hardware limits
  • Potential performance degradation
  • Use with caution

Too Tight

  • Model exceeds hardware capabilities
  • Will not run or perform very poorly
  • Not recommended for deployment

Technical Implementation

Automatic Quantization Selection

  • Starts with highest quality (least quantized) version
  • Progressively steps down through quantization levels
  • Stops when model fits within hardware constraints
  • Selects optimal balance of quality and compatibility

Backend Integration

Supports major inference engines out of the box:

Ollama

  • GGML format compatibility
  • Automatic model management
  • Cross-platform deployment

llama.cpp

  • Direct GGML model support
  • CPU-optimized inference
  • Extensive quantization options

MLX (Apple Silicon)

  • Native Metal compute acceleration
  • Unified memory optimization
  • Apple-specific optimizations

LM Studio

  • GUI-based model management
  • Multiple backend support
  • User-friendly deployment

Model Coverage

Comprehensive support for major model families:

  • Meta: Llama 2, Llama 3, Code Llama
  • Mistral: 7B, 22B, mixture of experts variants
  • Qwen: Qwen2.5, specialized variants
  • DeepSeek: Coder, Chat, and reasoning models
  • Hundreds of variants: Different sizes and quantizations

Usage Workflow

1. System Scanning

llmfit scan
  • Detects hardware specifications
  • Identifies available backends
  • Establishes performance baselines

2. Model Recommendation

llmfit recommend --use-case chat
  • Filters models by use case
  • Applies hardware constraints
  • Ranks by composite score

3. Deployment Guidance

  • Specific quantization recommendations
  • Backend selection advice
  • Memory usage predictions
  • Performance expectations

Output Format

Terminal Interface

  • Tabular display of compatible models
  • Color-coded compatibility indicators
  • Performance metrics (tok/s, memory %)
  • Quantization and backend recommendations

Key Metrics Displayed

  • Model name and parameter count
  • Composite compatibility score
  • Estimated tokens per second
  • Memory usage percentage
  • Recommended quantization level
  • Optimal backend/mode
  • Context window support
  • Fit status label

Benefits and Impact

Developer Experience

  • Eliminates trial-and-error: No more downloading incompatible models
  • Saves time: Quick identification of optimal models
  • Reduces frustration: Clear compatibility guidance
  • Improves success rate: Higher likelihood of successful deployment

Resource Efficiency

  • Bandwidth savings: Avoid downloading unusable models
  • Storage optimization: Only download compatible variants
  • Performance prediction: Set realistic expectations
  • Hardware utilization: Maximize available resources

Community Value

  • Open source: Free for all users
  • Extensible: Community contributions welcome
  • Educational: Teaches hardware-model relationships
  • Standardization: Common framework for model selection

Limitations and Considerations

Prediction Accuracy

  • Performance estimates based on heuristics
  • Actual performance may vary with specific workloads
  • Hardware-specific optimizations not fully captured

Model Coverage

  • Requires manual updates for new model releases
  • Emerging architectures may not be fully supported
  • Custom or fine-tuned models need separate evaluation

Environmental Factors

  • Doesn't account for thermal throttling
  • Background processes affect available resources
  • Dynamic hardware states not considered

Future Development

Enhanced Prediction

  • Machine learning-based performance modeling
  • Real-world benchmark integration
  • Dynamic hardware monitoring

Expanded Coverage

  • Broader model ecosystem support
  • Multi-modal model evaluation
  • Custom model analysis

Advanced Features

  • Cost-performance optimization
  • Multi-model deployment planning
  • Automated model updating

See also