← Alle LLMs
LLM · Final-Bench

Darwin

Eine Modellfamilie als lebendes Objekt: Varianten, ausgewählte Leistungsdaten und alle verbundenen BrunoSan-Quellensignale.

Erwähnungen
in 30 Tagen
Quellen
12Varianten
Im Universe erkunden →History öffnen →

Modelle

Alle bekannten Varianten dieser Familie. Ein direkter Artikel-Link springt exakt zur genannten Variante.

Final-Bench · Darwin

Darwin-180B-RSI

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
joelniklaus/LEXam-hardScore45.72Quelle ↗abgerufen 2026-10-10
Testdetails

LEXam-hard, all 518 open questions; single sample (temperature=1.0, top_p=0.95, top_k=20); thinking budget 32,768 tokens, responses truncated at 32K regenerated with a 120K budget (60 items); judged by DeepSeek-R1-0528 with the LEXam paper judge prompt (eval.yaml); score = mean of German and English mean grades x 100 (de 44.37, en 47.07); 1 item without a parsable grade counted as 0; bf16

Delores-Lin/MDPBenchScore83.65Quelle ↗abgerufen 2026-10-10
Testdetails

Official MDPBench code and prompt, public set 2,720 pages (text edit distance, formula CDM, table TEDS). vLLM FP8, temperature 0, thinking on, 32,768 max tokens; pages with an empty thinking-mode answer re-read with thinking off (17 pages). Single run.

MathArena/hmmt_feb_2026Score100Quelle ↗abgerufen 2026-10-10
Testdetails

HMMT February 2026, all 33 problems; majority vote over 16 samples (maj@16) = 100.0; mean accuracy over 16 samples = 96.59; temperature=1.0, top_p=0.95, top_k=20; thinking budget 131,072 tokens; bf16

MathArena/aime_2026Score100Quelle ↗abgerufen 2026-10-10
Testdetails

AIME 2026, all 30 problems; majority vote over 16 samples (maj@16) = 100.0; mean accuracy over 16 samples = 98.75; temperature=1.0, top_p=0.95, top_k=20; thinking budget 131,072 tokens; bf16

LEXam-Benchmark/LEXamScore68.94Quelle ↗abgerufen 2026-10-10
Testdetails

LEXam mcq_4_choices, all 1,655 items; majority vote over 4 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 32,768 tokens; bf16. Single-sample 60.54, mean of 4 samples 61.42

Weitere 4 Ergebnisse
BenchmarkWertBeleg
LiquidAI/ifstruct-v1.0Score98.95Quelle ↗abgerufen 2026-10-10
Testdetails

Official harness (github.com/Liquid4All/ifstruct, ifstruct-eval) with its defaults: temperature 0, max_tokens 16000, 32 threads, 2,000 prompts from data/test.jsonl. Thinking on. Served checkpoint: FINAL-Bench/Darwin-180B-RSI on vLLM 0.29.0 with online FP8 quantization, tensor parallel 2. 1,979 of 2,000 passed; one request exceeded the harness 120-second read timeout on every retry and is counted as a failure. Single run.

MMMU/MMMU_ProScore79.48Quelle ↗abgerufen 2026-10-10
Testdetails

MMMU-Pro vision setting, all 1,730 items; majority vote over 3 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16

Idavidrein/gpqaScore94.44Quelle ↗abgerufen 2026-10-10
Testdetails

GPQA Diamond, all 198 items; majority vote over up to 16 samples (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16

TIGER-Lab/MMLU-ProScore88.12Quelle ↗abgerufen 2026-10-10
Testdetails

MMLU-Pro test, all 12,032 items; single sample (temperature=1.0, top_p=0.95, top_k=20); thinking budget 131,072 tokens; bf16

Final-Bench · Darwin

Darwin-36B-Opus

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore88.4Quelle ↗abgerufen 2026-10-10
Testdetails

Standard inference, Pass@1

Final-Bench · Darwin

Darwin-4B-David

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore85Quelle ↗abgerufen 2026-10-10
Testdetails

4B class Gen-2 evolution, Pass@1

Final-Bench · Darwin

Darwin-397B-ZTC

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore93.43Quelle ↗abgerufen 2026-10-10
Testdetails

GPQA Diamond, all 198 items; greedy decoding (temperature=0), single sample (no voting / no test-time engine); weights released in FP8 (W8A8, compressed-tensors)

Final-Bench · Darwin

Darwin-31B-Opus

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore85.9Quelle ↗abgerufen 2026-10-10
Testdetails

Standard inference, Pass@1

Final-Bench · Darwin

Darwin-27B-Opus

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore86.9Quelle ↗abgerufen 2026-10-10
Testdetails

Standard inference, Pass@1

Final-Bench · Darwin

Darwin-60B-DUO

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore88.38Quelle ↗abgerufen 2026-10-10
Testdetails

DELPHI cascade system, Pass@1

Final-Bench · Darwin

Darwin-28B-REASON

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore89.39Quelle ↗abgerufen 2026-10-10
Testdetails

Darwin-DELPHI test-time engine, Pass@1

Final-Bench · Darwin

Darwin-398B-JGOS

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore90.9Quelle ↗abgerufen 2026-10-10
Testdetails

greedy decoding (temperature=0), single-sample (no voting / no test-time engine), max_tokens=16384, options shuffled seed=42; hardware: NVIDIA B200 x6 (TP2 x PP3), vLLM bfloat16

Final-Bench · Darwin

Darwin-9B-NEG

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
Idavidrein/gpqaScore84.34Quelle ↗abgerufen 2026-10-10
Testdetails

9B class, standard inference, Pass@1

Final-Bench · Darwin

Darwin-27B-ZTC

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
LocalLLaMA/typed-decisionsScore0.743Quelle ↗abgerufen 2026-10-10
Testdetails

General, zero-shot: never trained on the Typed Decisions train split, its workflows or its question schemas. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions (README request shape), all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options (scorer reproduces the README Uniform reference). One forward pass per question, no generated tokens.

Final-Bench · Darwin

Darwin-180B-RSI-R3

KNOWN

Technische Leistungsdaten sind für diese Variante noch nicht vollständig bestätigt.

—Kontext
—Max. Output
—Input
—Output
—Input / 1M
—Output / 1M
—Tools
—Reasoning
—Structured

Benchmarks

BenchmarkWertBeleg
llamaindex/ExtractBenchScore90.29Quelle ↗abgerufen 2026-10-10
Testdetails

Pipeline name: darwin_180b_rsi_r3_bf16_vllm_extract_oneshot_structured_output_file_nothink (vllm_extract provider, max_tokens 32768, temperature 0, json_object output, thinking disabled with chat_template_kwargs enable_thinking=false); served checkpoint: FINAL-Bench/Darwin-180B-RSI-R3 on vLLM 0.29.0 with online FP8 quantization, tensor parallel 2. Single run, 370 of 370 documents completed.

BrunoSan Chronik

Alle verbundenen Artikel, neueste zuerst.

Noch keine verbundenen Artikel.

Modelldaten und Nachrichten aus BrunoSan AI-News, OpenRouter und Hugging Face. Benchmark-Werte stammen aus den jeweils verlinkten Quellen; Testbedingungen können abweichen. Modelle bleiben auch ohne Benchmark-Werte sichtbar.