model-search-system-architecture.md

Model Search System Architecture

Overview

Model Search should become the place where a user can start from a vague goal:

  • "I need a fast local coding model for my Mac"
  • "I need a model I can fine-tune on this dataset"
  • "I need the cheapest cloud setup that still hits this benchmark target"

and quickly narrow millions of candidate repos and artifacts down to a small, justified shortlist.

This requires three systems working together:

  1. A model graph that understands lineage, formats, quantizations, benchmarks, and derivations.
  2. A hardware/runtime profiler that knows how different artifacts behave on real devices.
  3. A ranking engine that can combine use case, benchmark evidence, cost, latency, context, and memory constraints into a useful recommendation.

The goal is not just "index models". The goal is to answer:

  • Which model family is strongest for this task?
  • Which artifact is the best fit for this hardware?
  • What context window is realistic on this machine?
  • What prefill and decode speed should the user expect?
  • Is the right next step inference, quantization, fine-tuning, or dataset generation?

Direct Answer: Can We Estimate Theoretical TTFT / TPS?

Yes, but it should be treated as an estimator, not as a benchmark result.

The right approach is a hybrid model:

  • Analytical prior from model geometry, artifact format, quantization, runtime, and hardware profile.
  • Empirical calibration from measured local benchmarks on real hardware.

What can be estimated statically

These are good candidates for deterministic or near-deterministic estimation:

  • total parameter count
  • architecture class (dense, moe, hybrid)
  • active parameter count for supported MoE families
  • model weight footprint
  • KV-cache bytes per token
  • fast-only context ceiling
  • spill-into-slower-memory ceiling
  • rough prefill throughput region
  • rough decode throughput region

What should not be presented as ground truth without measurement

These should be labeled as estimates unless we have benchmark data:

  • prompt tokens/sec
  • generation tokens/sec
  • TTFT
  • energy draw in watts
  • joules per token

Product rule

Model Search should always distinguish:

  • Measured on this hardware
  • Estimated from profile
  • Inherited from base model / family prior

Those three signals are useful, but they are not interchangeable.


What Already Exists In bot0

The current codebase already has the beginnings of this system.

Catalog and metadata normalization

apps/bytespace/src/app/api/proxy/models/_shared.ts

This already aggregates and normalizes:

  • Hugging Face entries
  • OpenRouter entries
  • context window
  • local model geometry
  • artifact formats (gguf, mlx, transformers, safetensors)
  • some capability and category inference

Desktop model search UI

packages/desktop/src/components/Terminal/ExploreModelsPanel.tsx

This already renders:

  • model overview
  • context-window estimator
  • memory budgeting
  • local artifact fit information
  • provider performance data from OpenRouter

Local runtime loading

packages/desktop/electron/src/local-models.ts

This already manages:

  • GGUF loading
  • MLX loading
  • local context window launch parameters
  • runtime state

The next step is not to replace this. It is to formalize and extend it.


Core Design Principle

Model Search needs a canonical model graph, not just a flat search index.

Flat search is enough to find repos.

It is not enough to answer:

  • whether a repo is a fine-tune of another repo
  • whether a GGUF artifact and an MLX artifact come from the same base family
  • whether a benchmark result should be inherited, ignored, or treated as a weak prior
  • whether a quantized repo should be ranked as equivalent quality or only deployment-equivalent

That requires explicit nodes and edges.


The Model Graph

Node types

1. Canonical Model

The model family or canonical entity we want users to reason about.

Examples:

  • Qwen3.6 27B
  • Gemma 4 E2B
  • Mixtral 8x7B

Fields:

  • canonical id
  • display name
  • publisher
  • modality summary
  • architecture class
  • total params
  • active params
  • context window
  • benchmark rollup

2. Repo Variant

A specific upstream repo.

Examples:

  • Qwen/Qwen3.6-27B
  • mlx-community/Qwen3.6-27B-4bit
  • bartowski/Qwen3.6-27B-GGUF
  • unsloth/Qwen3.6-27B-GGUF

Fields:

  • repo id
  • source (huggingface, openrouter, core)
  • library (transformers, mlx, etc.)
  • model-card metadata
  • config metadata
  • file tree summary

3. Artifact

A runnable or downloadable deployment target.

Examples:

  • Q4_K_S GGUF
  • MLX 4-bit
  • bf16 safetensors
  • OpenRouter provider endpoint

Fields:

  • artifact id
  • format
  • quantization
  • size on disk
  • runtime family
  • estimated runtime footprint
  • KV precision assumptions
  • benchmark observations

4. Runtime Profile

The serving stack that executes an artifact.

Examples:

  • llama.cpp / Metal
  • mlx-lm
  • vLLM
  • SGLang
  • Transformers
  • OpenRouter provider endpoint

Fields:

  • runtime id
  • supported formats
  • KV cache type
  • offload policy
  • batching behavior
  • benchmark method

5. Hardware Profile

A normalized description of a target machine or cloud instance.

Examples:

  • Apple M4 / 34 GB unified / 22.91 GB fast GPU budget
  • 1x RTX 4090 / 24 GB
  • 8x H100 SXM
  • AWS g6e.2xlarge

Fields:

  • hardware id
  • vendor
  • accelerator type
  • memory topology
  • fast memory budget
  • total system memory
  • interconnect / spill bandwidth hints
  • peak compute metadata by dtype
  • benchmark observations
  • energy sensors available

6. Benchmark Observation

A measured result, not just a score claim.

Fields:

  • benchmark id
  • task
  • dataset
  • prompt shape
  • runtime
  • hardware
  • artifact
  • metric values
  • provenance

7. Dataset / Benchmark Definition

An official benchmark or a user-provided dataset.

Examples:

  • SWE-bench Verified
  • MMLU-Pro
  • GPQA
  • custom JSONL evaluation set

Graph edges

Important relationships:

  • DERIVES_FROM
  • BASE_MODEL_OF
  • QUANTIZED_AS
  • CONVERTED_TO
  • ADAPTER_OF
  • MERGED_FROM
  • BENCHMARKED_ON
  • RUNS_ON
  • BEST_FIT_FOR
  • CALIBRATED_ON
  • EVALUATED_WITH
  • FINE_TUNED_ON

This is what turns Model Search from a catalog into a decision engine.


Source Of Truth Hierarchy

Each field in Model Search should have an explicit priority order.

Parameter count

For Hugging Face, the priority order should be:

  1. model.safetensors.index.json metadata or safetensors metadata
  2. safetensors header parsing via HTTP range requests
  3. config-derived family logic
  4. model card text
  5. repo-name heuristic

This is the right direction because Hugging Face explicitly documents metadata parsing from safetensors without downloading the entire file.

Architecture class

Use config.json first.

Good signals include:

  • architectures
  • model_type
  • num_experts_per_tok
  • num_experts / num_local_experts
  • decoder_sparse_step
  • mlp_only_layers
  • layer_types

This should yield:

  • dense
  • moe
  • hybrid

Active parameters

Only show active params when we have:

  1. source-provided values, or
  2. a family-specific calculator we trust

If neither exists, leave it unknown.

Base model / lineage

Priority order:

  1. model card base_model
  2. model card base_model_relation
  3. adapter / merge metadata
  4. repo tags
  5. family-name heuristic

Benchmarks

Priority order:

  1. structured model-card eval results
  2. official benchmark leaderboard datasets
  3. bot0-owned benchmark runs
  4. family prior / inferred benchmark similarity

The last category must always be labeled as inferred, never as measured.


Performance Engine

The performance engine should be explicitly split into:

  1. memory fit
  2. speed estimate
  3. energy estimate
  4. measured observations

1. Memory fit

This part can be strongly deterministic.

Inputs:

  • model geometry
  • quantization
  • runtime
  • hardware profile
  • selected context window

Outputs:

  • download size
  • model runtime footprint
  • KV cache bytes at selected context
  • fast-only context max
  • spill range
  • unsafe range

This is already partly implemented in the current UI.

2. Speed estimate

This should use a roofline-style estimator plus runtime-specific correction factors.

TTFT decomposition

TTFT should be modeled as:

TTFT = queue + prompt processing + first decode step + transport/render overhead

For local inference, the important part is:

prompt processing + first decode step

Prefill estimate

Prefill should be estimated as:

prefill_time ≈ max(compute_bound_time, memory_bound_time) + runtime_overhead

Where:

  • compute_bound_time depends on prompt length, active parameters, architecture, and effective compute throughput for the runtime + quantization + hardware profile
  • memory_bound_time depends on bytes moved through weights, activations, and caches

Decode estimate

Decode should be estimated separately because it behaves differently:

decode_time_per_token ≈ max(weight_stream_time, kv_cache_time, compute_time) + sampling_overhead

For many local inference setups:

  • decode is often memory-bound
  • long context pushes more traffic through KV cache
  • once KV spills outside the fast memory budget, decode drops sharply

3. Energy estimate

Energy should use the same split:

  • predicted power band when measured sensors are unavailable
  • measured power / energy when counters are available

The best product metric is not just watts. It is:

  • average watts
  • peak watts
  • joules per prompt token
  • joules per generated token

That allows fair comparison across models and runtimes.

4. Measured observations

Measured results should override estimates where possible.

Examples:

  • prompt tok/sec
  • generation tok/sec
  • TTFT
  • GPU power draw
  • energy per token

These should be stored per:

  • artifact
  • runtime
  • hardware profile
  • prompt bucket
  • context bucket

Theoretical TPS / TTFT: What We Can Actually Build

Yes, we can build a theoretical estimator.

But it should be built as a two-layer model.

Layer 1: Analytical estimator

This estimator uses:

  • total params
  • active params
  • hidden size
  • layer count
  • KV head count
  • full-attention vs sliding-attention layers
  • quantization
  • runtime
  • prompt length
  • decode length
  • hardware memory bandwidth
  • hardware compute ceiling
  • fast memory size

This gives a first-principles estimate.

Layer 2: Empirical calibration

Then we calibrate the analytical estimate with benchmark observations.

This is necessary because raw hardware specs alone miss:

  • kernel quality
  • backend efficiency
  • batching policy
  • graph capture
  • prompt caching
  • runtime scheduler behavior
  • Metal vs CUDA vs HIP differences
  • vendor-specific quantization kernels

Product implication

We should not show a single naked number like:

  • 73 tok/sec

We should show:

  • Estimated decode: 58-74 tok/sec
  • Measured on similar M4 profile: 66 tok/sec
  • Confidence: medium

That is much more honest and much more useful.


KV cache

For decoder models, KV cache growth is linear in tokens and depends on:

  • number of growing cache layers
  • KV heads
  • K/V head dimensions
  • bytes per cache element

This part is deterministic enough to calculate directly from config-derived geometry.

Prefill

Prefill estimate should use:

  • prompt length bucket
  • active architecture path
  • runtime family
  • quantization class
  • hardware peak compute
  • hardware memory bandwidth

The result should be a throughput estimate in prompt tokens/sec and a TTFT estimate for a given prompt length.

Decode

Decode estimate should use:

  • selected context
  • active params
  • KV bytes/token
  • weight-stream size for the active path
  • spill amount beyond fast memory budget

The result should be:

  • fast-only decode estimate
  • spill decode estimate
  • one blended estimate for the current selection

Energy

Energy should be handled in three tiers:

  1. measured directly
  2. estimated from calibrated workload power profile
  3. rough fallback from power cap / TDP band

Tier 1 should always win.


Hardware Profiles

Each hardware profile should include both static specs and dynamic observability.

Static fields

  • vendor
  • chip / GPU name
  • total memory
  • fast memory budget
  • unified vs dedicated memory
  • theoretical memory bandwidth
  • theoretical compute ceilings by dtype
  • supported runtimes
  • sensor capabilities

Dynamic fields

  • measured prompt tok/sec by runtime
  • measured generation tok/sec by runtime
  • measured TTFT
  • measured watts
  • measured joules/token
  • thermal throttling behavior
  • spill penalties

Practical device APIs

Apple Silicon

Useful device information is already available from Metal:

  • hasUnifiedMemory
  • recommendedMaxWorkingSetSize
  • currentAllocatedSize

For energy, powermetrics can report estimated subsystem power, including GPU on supported systems, but Apple explicitly notes that the values are estimates and should not be used for cross-device comparisons.

NVIDIA

Useful sources:

  • nvidia-smi
  • NVML

These can expose:

  • instantaneous and average board power draw
  • memory power on supported devices
  • power limits

AMD

Useful sources:

  • AMD SMI / ROCm SMI

These can expose:

  • instant or average package power
  • energy counters
  • power caps

Design implication

Hardware profile collection should be pluggable by vendor, but normalized into one schema.


Benchmark Strategy

Benchmarks bot0 should ingest

1. Official benchmark leaderboards

Use Hugging Face benchmark datasets and leaderboard APIs as first-class sources.

This gives us:

  • official benchmark registry
  • structured benchmark scores
  • dataset ids
  • reproducible benchmark mapping

2. Model-card eval results

If model authors publish structured eval results in model cards, ingest them with provenance.

These are valuable but lower confidence than a bot0-owned rerun unless source quality is known.

3. bot0-owned local benchmarks

This is how Model Search becomes trustworthy for deployment advice.

GGUF

Use llama-bench and store:

  • prompt throughput (pp*)
  • generation throughput (tg*)
  • runtime metadata
MLX

Use mlx_lm.benchmark and store:

  • prompt tok/sec
  • generation tok/sec
  • memory footprint
Future cloud runtimes

For cloud mapping later:

  • vLLM
  • SGLang
  • provider-native endpoints

Benchmark buckets

Store measurements by:

  • model artifact
  • runtime
  • hardware profile
  • prompt length bucket
  • output length bucket
  • context bucket
  • batch size / concurrency

This is what makes TTFT and TPS predictions actually useful later.


Benchmark Inheritance And Family Priors

This is where the model graph matters.

Safe inheritance rules

Can be inherited strongly

  • architecture family
  • tokenizer family
  • context rules
  • modality support
  • base lineage

Can be inherited weakly

  • likely benchmark neighborhood
  • likely fine-tuning behavior
  • likely quantization sensitivity
  • likely memory/perf shape

Should not be inherited as fact

  • exact benchmark score
  • exact prompt tok/sec
  • exact generation tok/sec
  • exact energy draw

Product rule

If a fine-tune has no benchmark:

  • use base-model results as a prior
  • show a confidence penalty
  • label it as inferred

That still helps ranking without misleading the user.


User Flows This Enables

1. "Find me the best model for my hardware"

Inputs:

  • hardware profile
  • use case
  • latency target
  • minimum context

Outputs:

  • shortlist ranked by fit
  • which runtime to use
  • which quantization to use
  • what context is realistic
  • expected TTFT / decode band

2. "Find me the best model for my benchmark or dataset"

Inputs:

  • benchmark name or sample dataset
  • required modalities
  • cost / latency constraint

Outputs:

  • benchmark-backed candidates
  • confidence score
  • whether fine-tuning is likely needed

3. "Compare several models at once"

Inputs:

  • model set
  • hardware
  • benchmark or eval set

Outputs:

  • side-by-side speed
  • memory fit
  • benchmark quality
  • energy per token
  • cost per million tokens

4. "Should I fine-tune or just deploy a stronger base model?"

Inputs:

  • use case
  • examples
  • current candidate model

Outputs:

  • off-the-shelf candidates
  • expected lift from fine-tune
  • dataset needs
  • evaluation plan

Phase 1: Metadata hardening

Goal: make model identity and structure reliable.

Deliverables:

  • reliable total params from safetensors metadata
  • architecture classification from config
  • active params for supported MoE families
  • base model / merge / adapter graph
  • benchmark ingestion from model cards and benchmark datasets

Phase 2: Runtime benchmarking

Goal: make local performance trustworthy.

Deliverables:

  • GGUF benchmark runner
  • MLX benchmark runner
  • vendor power collectors
  • normalized benchmark result schema
  • hardware profile registry

Phase 3: Estimation engine

Goal: provide useful predictions before measurement exists.

Deliverables:

  • roofline-style estimator
  • spill penalty model
  • prompt/decode estimate bands
  • energy estimate bands
  • confidence scoring

Phase 4: Ranking engine

Goal: return real candidate lists, not just search results.

Deliverables:

  • use-case weighting
  • constraint solver
  • rank explanations
  • benchmark confidence weighting

Phase 5: Evaluation workspace

Goal: turn Model Search into an experimentation environment.

Deliverables:

  • run benchmark against one or more models
  • run dataset eval against one or more models
  • compare local and cloud deployment options
  • export fine-tuning / deployment recommendation

Labels to use

  • Measured on this machine
  • Measured on similar hardware
  • Estimated from profile
  • Inherited from base model
  • Structured benchmark result
  • Model-card reported

Labels to avoid

  • Fastest
  • Best
  • Guaranteed

unless the ranking objective is explicit.


Practical Recommendation For bot0

The right long-term architecture is:

  • graph-first for model lineage
  • benchmark-first for quality
  • profile-first for hardware fit
  • hybrid estimator + measurement for performance

That lets Model Search answer both immediate deployment questions and longer-term model strategy questions.

If the user only wants to run a local model, the system should answer:

  • what to download
  • which runtime to use
  • what context fits
  • expected fast-only behavior
  • expected spill behavior

If the user wants to choose a model family or plan a fine-tune, the system should answer:

  • what base families are strongest
  • what fine-tunes already exist
  • what benchmarks support them
  • what hardware is needed
  • whether the use case is best served by inference, fine-tuning, or dataset generation

That is the direction that turns Model Search into a real decision system rather than a prettier catalog.


References

Archived product

A chapter of bot0, preserved.

bot0 was a working product by Bytespace Labs. This site preserves its original design and product experience. The hosted service is no longer running; downloads, new accounts and purchases are unavailable.

Product descriptions, documentation and pricing reflect the product when it was active. The interactions preserved here are not connected to its former backend.

Interested in the technology?

We’re open to discussing an acquisition of the technology and codebase behind bot0.

Discuss an acquisition