Model Search System Architecture
Overview
Model Search should become the place where a user can start from a vague goal:
- "I need a fast local coding model for my Mac"
- "I need a model I can fine-tune on this dataset"
- "I need the cheapest cloud setup that still hits this benchmark target"
and quickly narrow millions of candidate repos and artifacts down to a small, justified shortlist.
This requires three systems working together:
- A model graph that understands lineage, formats, quantizations, benchmarks, and derivations.
- A hardware/runtime profiler that knows how different artifacts behave on real devices.
- A ranking engine that can combine use case, benchmark evidence, cost, latency, context, and memory constraints into a useful recommendation.
The goal is not just "index models". The goal is to answer:
- Which model family is strongest for this task?
- Which artifact is the best fit for this hardware?
- What context window is realistic on this machine?
- What prefill and decode speed should the user expect?
- Is the right next step inference, quantization, fine-tuning, or dataset generation?
Direct Answer: Can We Estimate Theoretical TTFT / TPS?
Yes, but it should be treated as an estimator, not as a benchmark result.
The right approach is a hybrid model:
- Analytical prior from model geometry, artifact format, quantization, runtime, and hardware profile.
- Empirical calibration from measured local benchmarks on real hardware.
What can be estimated statically
These are good candidates for deterministic or near-deterministic estimation:
- total parameter count
- architecture class (
dense,moe,hybrid) - active parameter count for supported MoE families
- model weight footprint
- KV-cache bytes per token
- fast-only context ceiling
- spill-into-slower-memory ceiling
- rough prefill throughput region
- rough decode throughput region
What should not be presented as ground truth without measurement
These should be labeled as estimates unless we have benchmark data:
- prompt tokens/sec
- generation tokens/sec
- TTFT
- energy draw in watts
- joules per token
Product rule
Model Search should always distinguish:
- Measured on this hardware
- Estimated from profile
- Inherited from base model / family prior
Those three signals are useful, but they are not interchangeable.
What Already Exists In bot0
The current codebase already has the beginnings of this system.
Catalog and metadata normalization
apps/bytespace/src/app/api/proxy/models/_shared.ts
This already aggregates and normalizes:
- Hugging Face entries
- OpenRouter entries
- context window
- local model geometry
- artifact formats (
gguf,mlx,transformers,safetensors) - some capability and category inference
Desktop model search UI
packages/desktop/src/components/Terminal/ExploreModelsPanel.tsx
This already renders:
- model overview
- context-window estimator
- memory budgeting
- local artifact fit information
- provider performance data from OpenRouter
Local runtime loading
packages/desktop/electron/src/local-models.ts
This already manages:
- GGUF loading
- MLX loading
- local context window launch parameters
- runtime state
The next step is not to replace this. It is to formalize and extend it.
Core Design Principle
Model Search needs a canonical model graph, not just a flat search index.
Flat search is enough to find repos.
It is not enough to answer:
- whether a repo is a fine-tune of another repo
- whether a GGUF artifact and an MLX artifact come from the same base family
- whether a benchmark result should be inherited, ignored, or treated as a weak prior
- whether a quantized repo should be ranked as equivalent quality or only deployment-equivalent
That requires explicit nodes and edges.
The Model Graph
Node types
1. Canonical Model
The model family or canonical entity we want users to reason about.
Examples:
Qwen3.6 27BGemma 4 E2BMixtral 8x7B
Fields:
- canonical id
- display name
- publisher
- modality summary
- architecture class
- total params
- active params
- context window
- benchmark rollup
2. Repo Variant
A specific upstream repo.
Examples:
Qwen/Qwen3.6-27Bmlx-community/Qwen3.6-27B-4bitbartowski/Qwen3.6-27B-GGUFunsloth/Qwen3.6-27B-GGUF
Fields:
- repo id
- source (
huggingface,openrouter,core) - library (
transformers,mlx, etc.) - model-card metadata
- config metadata
- file tree summary
3. Artifact
A runnable or downloadable deployment target.
Examples:
Q4_K_S GGUFMLX 4-bitbf16 safetensorsOpenRouter provider endpoint
Fields:
- artifact id
- format
- quantization
- size on disk
- runtime family
- estimated runtime footprint
- KV precision assumptions
- benchmark observations
4. Runtime Profile
The serving stack that executes an artifact.
Examples:
llama.cpp / Metalmlx-lmvLLMSGLangTransformersOpenRouter provider endpoint
Fields:
- runtime id
- supported formats
- KV cache type
- offload policy
- batching behavior
- benchmark method
5. Hardware Profile
A normalized description of a target machine or cloud instance.
Examples:
Apple M4 / 34 GB unified / 22.91 GB fast GPU budget1x RTX 4090 / 24 GB8x H100 SXMAWS g6e.2xlarge
Fields:
- hardware id
- vendor
- accelerator type
- memory topology
- fast memory budget
- total system memory
- interconnect / spill bandwidth hints
- peak compute metadata by dtype
- benchmark observations
- energy sensors available
6. Benchmark Observation
A measured result, not just a score claim.
Fields:
- benchmark id
- task
- dataset
- prompt shape
- runtime
- hardware
- artifact
- metric values
- provenance
7. Dataset / Benchmark Definition
An official benchmark or a user-provided dataset.
Examples:
SWE-bench VerifiedMMLU-ProGPQA- custom JSONL evaluation set
Graph edges
Important relationships:
DERIVES_FROMBASE_MODEL_OFQUANTIZED_ASCONVERTED_TOADAPTER_OFMERGED_FROMBENCHMARKED_ONRUNS_ONBEST_FIT_FORCALIBRATED_ONEVALUATED_WITHFINE_TUNED_ON
This is what turns Model Search from a catalog into a decision engine.
Source Of Truth Hierarchy
Each field in Model Search should have an explicit priority order.
Parameter count
For Hugging Face, the priority order should be:
model.safetensors.index.jsonmetadata or safetensors metadata- safetensors header parsing via HTTP range requests
- config-derived family logic
- model card text
- repo-name heuristic
This is the right direction because Hugging Face explicitly documents metadata parsing from safetensors without downloading the entire file.
Architecture class
Use config.json first.
Good signals include:
architecturesmodel_typenum_experts_per_toknum_experts/num_local_expertsdecoder_sparse_stepmlp_only_layerslayer_types
This should yield:
densemoehybrid
Active parameters
Only show active params when we have:
- source-provided values, or
- a family-specific calculator we trust
If neither exists, leave it unknown.
Base model / lineage
Priority order:
- model card
base_model - model card
base_model_relation - adapter / merge metadata
- repo tags
- family-name heuristic
Benchmarks
Priority order:
- structured model-card eval results
- official benchmark leaderboard datasets
- bot0-owned benchmark runs
- family prior / inferred benchmark similarity
The last category must always be labeled as inferred, never as measured.
Performance Engine
The performance engine should be explicitly split into:
- memory fit
- speed estimate
- energy estimate
- measured observations
1. Memory fit
This part can be strongly deterministic.
Inputs:
- model geometry
- quantization
- runtime
- hardware profile
- selected context window
Outputs:
- download size
- model runtime footprint
- KV cache bytes at selected context
- fast-only context max
- spill range
- unsafe range
This is already partly implemented in the current UI.
2. Speed estimate
This should use a roofline-style estimator plus runtime-specific correction factors.
TTFT decomposition
TTFT should be modeled as:
TTFT = queue + prompt processing + first decode step + transport/render overhead
For local inference, the important part is:
prompt processing + first decode step
Prefill estimate
Prefill should be estimated as:
prefill_time ≈ max(compute_bound_time, memory_bound_time) + runtime_overhead
Where:
compute_bound_timedepends on prompt length, active parameters, architecture, and effective compute throughput for the runtime + quantization + hardware profilememory_bound_timedepends on bytes moved through weights, activations, and caches
Decode estimate
Decode should be estimated separately because it behaves differently:
decode_time_per_token ≈ max(weight_stream_time, kv_cache_time, compute_time) + sampling_overhead
For many local inference setups:
- decode is often memory-bound
- long context pushes more traffic through KV cache
- once KV spills outside the fast memory budget, decode drops sharply
3. Energy estimate
Energy should use the same split:
- predicted power band when measured sensors are unavailable
- measured power / energy when counters are available
The best product metric is not just watts. It is:
- average watts
- peak watts
- joules per prompt token
- joules per generated token
That allows fair comparison across models and runtimes.
4. Measured observations
Measured results should override estimates where possible.
Examples:
prompt tok/secgeneration tok/secTTFTGPU power drawenergy per token
These should be stored per:
- artifact
- runtime
- hardware profile
- prompt bucket
- context bucket
Theoretical TPS / TTFT: What We Can Actually Build
Yes, we can build a theoretical estimator.
But it should be built as a two-layer model.
Layer 1: Analytical estimator
This estimator uses:
- total params
- active params
- hidden size
- layer count
- KV head count
- full-attention vs sliding-attention layers
- quantization
- runtime
- prompt length
- decode length
- hardware memory bandwidth
- hardware compute ceiling
- fast memory size
This gives a first-principles estimate.
Layer 2: Empirical calibration
Then we calibrate the analytical estimate with benchmark observations.
This is necessary because raw hardware specs alone miss:
- kernel quality
- backend efficiency
- batching policy
- graph capture
- prompt caching
- runtime scheduler behavior
- Metal vs CUDA vs HIP differences
- vendor-specific quantization kernels
Product implication
We should not show a single naked number like:
73 tok/sec
We should show:
Estimated decode: 58-74 tok/secMeasured on similar M4 profile: 66 tok/secConfidence: medium
That is much more honest and much more useful.
Recommended Math Model
KV cache
For decoder models, KV cache growth is linear in tokens and depends on:
- number of growing cache layers
- KV heads
- K/V head dimensions
- bytes per cache element
This part is deterministic enough to calculate directly from config-derived geometry.
Prefill
Prefill estimate should use:
- prompt length bucket
- active architecture path
- runtime family
- quantization class
- hardware peak compute
- hardware memory bandwidth
The result should be a throughput estimate in prompt tokens/sec and a TTFT estimate for a given prompt length.
Decode
Decode estimate should use:
- selected context
- active params
- KV bytes/token
- weight-stream size for the active path
- spill amount beyond fast memory budget
The result should be:
- fast-only decode estimate
- spill decode estimate
- one blended estimate for the current selection
Energy
Energy should be handled in three tiers:
- measured directly
- estimated from calibrated workload power profile
- rough fallback from power cap / TDP band
Tier 1 should always win.
Hardware Profiles
Each hardware profile should include both static specs and dynamic observability.
Static fields
- vendor
- chip / GPU name
- total memory
- fast memory budget
- unified vs dedicated memory
- theoretical memory bandwidth
- theoretical compute ceilings by dtype
- supported runtimes
- sensor capabilities
Dynamic fields
- measured prompt tok/sec by runtime
- measured generation tok/sec by runtime
- measured TTFT
- measured watts
- measured joules/token
- thermal throttling behavior
- spill penalties
Practical device APIs
Apple Silicon
Useful device information is already available from Metal:
hasUnifiedMemoryrecommendedMaxWorkingSetSizecurrentAllocatedSize
For energy, powermetrics can report estimated subsystem power, including GPU on supported systems, but Apple explicitly notes that the values are estimates and should not be used for cross-device comparisons.
NVIDIA
Useful sources:
nvidia-smi- NVML
These can expose:
- instantaneous and average board power draw
- memory power on supported devices
- power limits
AMD
Useful sources:
- AMD SMI / ROCm SMI
These can expose:
- instant or average package power
- energy counters
- power caps
Design implication
Hardware profile collection should be pluggable by vendor, but normalized into one schema.
Benchmark Strategy
Benchmarks bot0 should ingest
1. Official benchmark leaderboards
Use Hugging Face benchmark datasets and leaderboard APIs as first-class sources.
This gives us:
- official benchmark registry
- structured benchmark scores
- dataset ids
- reproducible benchmark mapping
2. Model-card eval results
If model authors publish structured eval results in model cards, ingest them with provenance.
These are valuable but lower confidence than a bot0-owned rerun unless source quality is known.
3. bot0-owned local benchmarks
This is how Model Search becomes trustworthy for deployment advice.
GGUF
Use llama-bench and store:
- prompt throughput (
pp*) - generation throughput (
tg*) - runtime metadata
MLX
Use mlx_lm.benchmark and store:
- prompt tok/sec
- generation tok/sec
- memory footprint
Future cloud runtimes
For cloud mapping later:
vLLMSGLang- provider-native endpoints
Benchmark buckets
Store measurements by:
- model artifact
- runtime
- hardware profile
- prompt length bucket
- output length bucket
- context bucket
- batch size / concurrency
This is what makes TTFT and TPS predictions actually useful later.
Benchmark Inheritance And Family Priors
This is where the model graph matters.
Safe inheritance rules
Can be inherited strongly
- architecture family
- tokenizer family
- context rules
- modality support
- base lineage
Can be inherited weakly
- likely benchmark neighborhood
- likely fine-tuning behavior
- likely quantization sensitivity
- likely memory/perf shape
Should not be inherited as fact
- exact benchmark score
- exact prompt tok/sec
- exact generation tok/sec
- exact energy draw
Product rule
If a fine-tune has no benchmark:
- use base-model results as a prior
- show a confidence penalty
- label it as inferred
That still helps ranking without misleading the user.
User Flows This Enables
1. "Find me the best model for my hardware"
Inputs:
- hardware profile
- use case
- latency target
- minimum context
Outputs:
- shortlist ranked by fit
- which runtime to use
- which quantization to use
- what context is realistic
- expected TTFT / decode band
2. "Find me the best model for my benchmark or dataset"
Inputs:
- benchmark name or sample dataset
- required modalities
- cost / latency constraint
Outputs:
- benchmark-backed candidates
- confidence score
- whether fine-tuning is likely needed
3. "Compare several models at once"
Inputs:
- model set
- hardware
- benchmark or eval set
Outputs:
- side-by-side speed
- memory fit
- benchmark quality
- energy per token
- cost per million tokens
4. "Should I fine-tune or just deploy a stronger base model?"
Inputs:
- use case
- examples
- current candidate model
Outputs:
- off-the-shelf candidates
- expected lift from fine-tune
- dataset needs
- evaluation plan
Recommended Rollout Plan
Phase 1: Metadata hardening
Goal: make model identity and structure reliable.
Deliverables:
- reliable total params from safetensors metadata
- architecture classification from config
- active params for supported MoE families
- base model / merge / adapter graph
- benchmark ingestion from model cards and benchmark datasets
Phase 2: Runtime benchmarking
Goal: make local performance trustworthy.
Deliverables:
- GGUF benchmark runner
- MLX benchmark runner
- vendor power collectors
- normalized benchmark result schema
- hardware profile registry
Phase 3: Estimation engine
Goal: provide useful predictions before measurement exists.
Deliverables:
- roofline-style estimator
- spill penalty model
- prompt/decode estimate bands
- energy estimate bands
- confidence scoring
Phase 4: Ranking engine
Goal: return real candidate lists, not just search results.
Deliverables:
- use-case weighting
- constraint solver
- rank explanations
- benchmark confidence weighting
Phase 5: Evaluation workspace
Goal: turn Model Search into an experimentation environment.
Deliverables:
- run benchmark against one or more models
- run dataset eval against one or more models
- compare local and cloud deployment options
- export fine-tuning / deployment recommendation
Recommended Product Semantics
Labels to use
Measured on this machineMeasured on similar hardwareEstimated from profileInherited from base modelStructured benchmark resultModel-card reported
Labels to avoid
FastestBestGuaranteed
unless the ranking objective is explicit.
Practical Recommendation For bot0
The right long-term architecture is:
- graph-first for model lineage
- benchmark-first for quality
- profile-first for hardware fit
- hybrid estimator + measurement for performance
That lets Model Search answer both immediate deployment questions and longer-term model strategy questions.
If the user only wants to run a local model, the system should answer:
- what to download
- which runtime to use
- what context fits
- expected fast-only behavior
- expected spill behavior
If the user wants to choose a model family or plan a fine-tune, the system should answer:
- what base families are strongest
- what fine-tunes already exist
- what benchmarks support them
- what hardware is needed
- whether the use case is best served by inference, fine-tuning, or dataset generation
That is the direction that turns Model Search into a real decision system rather than a prettier catalog.
References
- Hugging Face Safetensors metadata parsing:
- Hugging Face model cards and base model metadata:
- Hugging Face benchmark leaderboard data:
- Hugging Face KV cache docs:
- MoE config examples:
- GGUF benchmarking:
- MLX benchmarking:
- OpenRouter endpoints and routing metrics:
- Apple Metal device memory inspection:
- Apple
powermetrics:man powermetrics
- NVIDIA power telemetry:
- AMD power telemetry: