Source first
A field starts with observable evidence from the listed source repository or linked publisher documentation — not from naming conventions.
OPEN MODEL WEIGHTS · EVIDENCE METHODOLOGY
The registry is designed to maximize useful coverage without converting assumptions into facts. Every important field should resolve to source evidence, an explicit classification rule, a labeled derivation — or an honest unknown.
CORE PRINCIPLES
The goal is not to make every field look complete. The goal is to make every published claim inspectable.
A field starts with observable evidence from the listed source repository or linked publisher documentation — not from naming conventions.
Verification applies to individual claims. A record can contain verified fields alongside explicit unknown or not-disclosed values.
Missing evidence is preserved as missing. Open Model Weights does not fill gaps just to make records look complete.
Repository revisions and verified-field changes become observed evidence over time instead of being overwritten by the newest state.
REGISTRY LIFECYCLE
Discovery, verification, revision checks and history are separate stages.
Identify a source repository that appears to publish open-weight model artifacts.
Require recognized weight artifacts and enough source evidence to create a real model record.
Check repository/API/config/model-card/license evidence field by field.
Store verified, declared, derived, not-disclosed or unknown states instead of silently inferring.
Read the current repository revision and fresh API metadata on the daily registry run.
When source evidence changes, create new observed history and field-level diffs.
FIELD → EVIDENCE → RULE
Verification is specific to the field. One source does not automatically validate the whole record.
| Field group | Primary evidence | Method / boundary |
|---|---|---|
| Identity & weights | Source repository API + exact file listing | Repository exists, recognized weight artifacts are present, exact filenames are retained. |
| License | Model-card metadata + repository license files / linked terms | Declared terms are recorded with evidence. Commercial-use classification is a comparison aid, not legal advice. |
| Context & parameters | Structured config/API first; explicit source text second | Structured values are preferred. Naming conventions do not become facts. |
| Formats & precision | Observed repository artifacts, filenames and dtype/config signals | Positive signals are recorded. “Not observed” does not mean no third-party conversion exists. |
| Lineage | Declared base-model metadata | No parent model is invented from similarity, architecture family or naming. |
| Training assets | Repository files + obvious model-card disclosure | Disclosure signals are recorded; Open Model Weights does not reconstruct undisclosed training. |
| Runtime support | Source tags, model card and repository artifacts | Compatibility is source-derived unless a runtime is explicitly labeled as independently tested. |
| Hardware memory | Derived from verified parameter count | Weight-only estimate: excludes KV cache, activations, runtime overhead and sharding. |
| Popularity & freshness | Source repository API metadata | Downloads/likes aid discovery, not quality ranking. Revision checks are distinct from full field verification. |
EVIDENCE STATES
Open Model Weights preserves the difference between something we directly checked, something the source merely declares, something we calculate, and something the available evidence does not establish.
Directly checked against the named evidence source.
Present in source metadata or publisher text, but semantically a publisher/source declaration.
Calculated from verified inputs and labeled as a derivation.
The checked standard evidence did not expose the value.
Available evidence is insufficient for a defensible value.
DERIVED VALUES
They estimate storage for model weights only — not end-to-end deployment memory.
Approximate weight-only memory for 16-bit weights.
Approximate weight-only memory for 8-bit weights.
Approximate weight-only memory for 4-bit weights.
KV cache, activations, optimizer state, runtime overhead, quantization metadata, device placement and sharding are not included. A displayed memory estimate is therefore not a deployment guarantee.
FRESHNESS & DATES
The daily pipeline checks current repository revision and fresh API metadata. If a source revision changes, relevant evidence is fetched again for field verification. The last revision check and the last full field verification are retained as distinct signals.
Used to detect source changes without re-fetching every unchanged artifact.
The most recent run that re-evaluated the relevant record fields.
Used only when no separate structured release date is available, and labeled as a proxy.
A changed timestamp does not automatically become a semantic model release.
BOUNDARIES
These limits are part of the methodology, not footnotes. They prevent useful discovery signals from being overstated as stronger evidence.
A blank field is preferable to a plausible but unsupported value.
The indexed Hugging Face URL is called the source repository unless publisher ownership is independently established.
Weight-memory estimates are comparison aids, not claims that a model will run within that amount of RAM or VRAM.
Runtime entries remain source-derived unless an execution test is explicitly documented.
Commercial-use labels summarize checked terms for comparison and are not legal advice.
Downloads and likes remain discovery signals and never become a model-quality ranking.
REPRODUCIBILITY
The methodology is backed by public data surfaces rather than a closed scoring system.
Report the model, the field in question and the strongest source you have. Public version control keeps changes attributable and inspectable.