Darwin-180B-RSI
KNOWNTechnische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.
Benchmarks
| Benchmark | Wert | Beleg |
|---|---|---|
| joelniklaus/LEXam-hardScore | 45.72 | Quelle ↗abgerufen 2026-10-10TestdetailsLEXam-hard, all 518 open questions; single sample (temperature=1.0, top_p=0.95, top_k=20); thinking budget 32,768 tokens, responses truncated at 32K regenerated with a 120K budget (60 items); judged by DeepSeek-R1-0528 with the LEXam paper judge prompt (eval.yaml); score = mean of German and English mean grades x 100 (de 44.37, en 47.07); 1 item without a parsable grade counted as 0; bf16 |
| Delores-Lin/MDPBenchScore | 83.65 | Quelle ↗abgerufen 2026-10-10TestdetailsOfficial MDPBench code and prompt, public set 2,720 pages (text edit distance, formula CDM, table TEDS). vLLM FP8, temperature 0, thinking on, 32,768 max tokens; pages with an empty thinking-mode answer re-read with thinking off (17 pages). Single run. |
| MathArena/hmmt_feb_2026Score | 100 | Quelle ↗abgerufen 2026-10-10TestdetailsHMMT February 2026, all 33 problems; majority vote over 16 samples (maj@16) = 100.0; mean accuracy over 16 samples = 96.59; temperature=1.0, top_p=0.95, top_k=20; thinking budget 131,072 tokens; bf16 |
| MathArena/aime_2026Score | 100 | Quelle ↗abgerufen 2026-10-10TestdetailsAIME 2026, all 30 problems; majority vote over 16 samples (maj@16) = 100.0; mean accuracy over 16 samples = 98.75; temperature=1.0, top_p=0.95, top_k=20; thinking budget 131,072 tokens; bf16 |
| LEXam-Benchmark/LEXamScore | 68.94 | Quelle ↗abgerufen 2026-10-10TestdetailsLEXam mcq_4_choices, all 1,655 items; majority vote over 4 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 32,768 tokens; bf16. Single-sample 60.54, mean of 4 samples 61.42 |
Weitere 4 Ergebnisse
| Benchmark | Wert | Beleg |
|---|---|---|
| LiquidAI/ifstruct-v1.0Score | 98.95 | Quelle ↗abgerufen 2026-10-10TestdetailsOfficial harness (github.com/Liquid4All/ifstruct, ifstruct-eval) with its defaults: temperature 0, max_tokens 16000, 32 threads, 2,000 prompts from data/test.jsonl. Thinking on. Served checkpoint: FINAL-Bench/Darwin-180B-RSI on vLLM 0.29.0 with online FP8 quantization, tensor parallel 2. 1,979 of 2,000 passed; one request exceeded the harness 120-second read timeout on every retry and is counted as a failure. Single run. |
| MMMU/MMMU_ProScore | 79.48 | Quelle ↗abgerufen 2026-10-10TestdetailsMMMU-Pro vision setting, all 1,730 items; majority vote over 3 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16 |
| Idavidrein/gpqaScore | 94.44 | Quelle ↗abgerufen 2026-10-10TestdetailsGPQA Diamond, all 198 items; majority vote over up to 16 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16 |
| TIGER-Lab/MMLU-ProScore | 88.12 | Quelle ↗abgerufen 2026-10-10TestdetailsMMLU-Pro test, all 12,032 items; single sample (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16 |