Model release

Darwin-27B-ZTC publishes open-weight, single-pass judging with corrected benchmark slices

FINAL-Bench releases a Safetensors classifier that produces typed decision probabilities without generative decoding; its creators have corrected initially inconsistent per-type benchmark figures.

Open Model Weights published9 Oct 2026
Primary sourceFINAL-Bench / Hugging Face model card
Source published2026-10-08

FINAL-Bench has published Darwin-27B-ZTC, an open-weight decision model intended for evaluating answers and routing automated workflows. Rather than generating a written judgment token by token, the model processes a state and a typed question, then returns probabilities for allowed outcomes. The publisher describes three question formats: free-form correctness, selection among choices and scoring against an ordered rubric. The design may simplify downstream systems that need a probability distribution rather than another piece of model-generated prose.

The official Hugging Face repository lists an Apache-2.0 license and downloadable BF16 Safetensors backbone weights, along with a separate readout head, decision configuration and inference code. The model name says 27B, while the Hub metadata lists approximately 26B parameters; these are distinct publisher-facing representations that should not be silently collapsed into a newly verified count. The model card says running the checkpoint requires GPU memory for roughly 54 GB of BF16 weights. That is a deployment note, not an Open Model Weights hardware measurement.

FINAL-Bench reports 0.743 overall accuracy over 2,000 judgments in a zero-shot typed-decisions evaluation, with a Brier score of 0.097 and KL divergence of 0.204. These are publisher-reported benchmark values, not results independently replicated by Open Model Weights. An important evidence-quality detail is already visible: a community reviewer found that the original per-type results did not reconcile with the headline accuracy. The publisher acknowledged mixing numbers from two runs and corrected the official model card. Its updated per-type values are 0.845 for binary correctness, 0.732 for multiple choice and 0.675 for scoring.

A single forward pass avoids autoregressive output decoding but does not eliminate input processing or hardware costs. The authors also caution implicitly through their evaluation setup: calibration on one benchmark does not establish reliability on another organization's production data. Before recommending thresholds for automatic acceptance, users should measure local error rates by category, verify that reported probabilities are calibrated, and inspect changes across model revisions. For the OMW registry, this announcement creates a source-linked candidate for field verification, not a published OMW benchmark or a verified change to any existing model.

OMW REGISTRY WATCH

What this changes for the evidence layer

New model candidate, not a verified OMW registry record. Pin the publisher repository SHA, confirm the license file and exact BF16 Safetensors shards, readout head, config and inference implementation. Treat the reported accuracy, calibration and hardware needs as publisher claims, not OMW measurements. Preserve the public correction to the first benchmark per-type table and investigate any future changed revisions.

PRIMARY SOURCE

FINAL-Bench / Hugging Face model card

This brief is based on the cited primary source. Performance, benchmark and comparative claims remain attributed unless Open Model Weights publishes an independent measurement.

Open source article ↗