aplomb-1
KNOWNTechnische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.
Benchmarks
| Benchmark | Wert | Beleg |
|---|---|---|
| LocalLLaMA/typed-decisionsScore | 0.737 | Quelle ↗abgerufen 2026-10-10TestdetailsGeneral model, not zero-shot on this benchmark: its training data included the train split (all 1,200 cases). The test split was not used for training; it was one of the suites used to choose between checkpoints. Scored on all 400 test cases (2,000 decisions, no errors) through the hosted EmpirioLabs API, model aplomb-1, one request per case with all five questions, on 2026-10-07. Accuracy: argmax of the prediction against the argmax of the gold distribution, per question, averaged over the 20 questions. |