Phocinae-Largha-150M-v1
KNOWNTechnische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.
Benchmarks
| Benchmark | Wert | Beleg |
|---|---|---|
| LocalLLaMA/typed-decisionsScore | 0.906 | Quelle ↗abgerufen 2026-10-10TestdetailsSpecialist: fitted on the Typed Decisions train split (mmBERT-small base) — NOT zero-shot; per the benchmark README this belongs to the "fitted/fine-tuned on train" table and is not comparable with the zero-shot table. Test split, 400 cases / 2,000 decisions, one request per case with the state and all five questions (README request shape), all answered, zero errors. Accuracy = agreement with the gold label. One forward pass per question, no generated tokens. Self-reported (self-host, no verifyToken). The submitted accuracy sits above the 0.735 teacher self-agreement ceiling — per the README this indicates the model is learning the teacher's quirks (label-specificity), also probed via JevBench on the model card. zh (machine-translated test cases; the training mix includes machine-translated + native Chinese rows) has no benchmark task; see the model card. |