VERTICAL::ARXIV · LIVE

VERTICAL::ARXIV · LIVE RESEARCH INTELLIGENCE

Know which research matters.
Connect the evidence.

Search 606,189 AI papers as connected objects—authors, entities, citations, repositories and affiliation signals—through one traceable research graph.

606.2KPapers indexedSource coverage 2007-04-01 → 2026-07-30
11.7MReferences parsedReference records extracted from paper HTML
2.4MResolved citation edgesReferences resolving to addressable ArXiv papers
136.1KPapers linked to repositoriesDistinct papers with at least one canonical repository relation
144.9KCanonical repositoriesLowercase canonical GitHub owner/repository identities
484.3KEnriched empirical papersIdentified inside 529,596 enriched papers

SNAPSHOT::2026-08-09T13:45:06.005587Z · COUNTS::PRODUCTION_DB · STATUS_FRONTIER::2007-03-07T00:00:00Z

DECISION VALUE

Questions normal search
cannot answer cleanly.

The value is not another list of papers. It is the ability to query relationships, influence, code and research trajectories as explicit objects.

QUESTION::INFLUENCE

Which papers are shaping the field?

Rank papers by inbound citation edges that resolve inside the live ArXiv graph, then inspect who cites them.

arxiv_most_cited → arxiv_citation_network
QUESTION::INTERSECTION

Where do two ideas meet?

Resolve the exact paper set connecting two methods, models, datasets or benchmarks instead of guessing from keywords.

arxiv_co_occurrence(entity_a, entity_b)
QUESTION::CODE

Which research ships usable code?

Filter enriched papers for code and map the wider graph to canonical GitHub repositories and organisations.

arxiv_search_papers(has_code_only=True) → arxiv_repo_landscape

OBJECT MODEL

Documents become
an inspectable graph.

Every layer contributes a different form of evidence. The MCP interface exposes the same relations your team sees here.

OBJECT::PAPERPaperArXiv ID, title, abstract, category, date
OBJECT::AUTHORAuthorPosition and connected paper portfolio
OBJECT::ENTITYEntityMethod, model, benchmark or dataset
EDGE::CITATIONCitationInbound and outbound resolved relations
OBJECT::REPOSITORYRepositoryCanonical GitHub owner/repository identity
SIGNAL::AFFILIATIONInstitutionHTML-extracted affiliation and organisation signals

One paper object can resolve its ordered authors, extracted entities, reference records, incoming citations and canonical repositories without losing the source identity that produced each relation.

LIVE CAPABILITIES

Already queryable.
Not “coming soon.”

13 public tools are derived directly from the deployed MCP source registry. The page does not maintain a separate tool count.

CAPABILITY::CITATION_GRAPH

Citation networks

Traverse incoming and outgoing ArXiv citation relations from one paper.

TOOL::arxiv_citation_network
CAPABILITY::ENTITY_TRENDS

Research trajectories

Measure how methods, models, benchmarks and datasets rise or decline over time.

TOOL::arxiv_entity_trend
CAPABILITY::CO_OCCURRENCE

Concept intersections

Find the exact papers that connect two named research entities.

TOOL::arxiv_co_occurrence
CAPABILITY::CODE_GRAPH

Research code

Map papers to canonical repositories and the organisations publishing them.

TOOL::arxiv_repo_landscape
CAPABILITY::INSTITUTIONS

Institution signals

Compare affiliation and GitHub-organisation signals with explicit coverage caveats.

TOOL::arxiv_institution_ranking
CAPABILITY::RELATED

Cross-domain intelligence

Connect research to AI News, Robotics, Quantum and Biotech signals.

TOOL::get_related_intelligence

LIVE RESEARCH SIGNALS

See what the graph
knows right now.

Entity lists below are ordered by title mentions, a stricter relevance signal than incidental mentions anywhere in the paper.

TITLE FOCUS · METHODS

  1. #01Transformer9.2K
  2. #02Multi-Agent5.8K
  3. #03Diffusion5.7K
  4. #04Knowledge Distillation4.2K
  5. #05Zero-Shot4.1K
  6. #06Few-Shot3.8K
  7. #07VLM3.2K
  8. #08GAN3.0K

TITLE FOCUS · MODELS

  1. #01BERT1.4K
  2. #02CLIP1.0K
  3. #03ChatGPT913
  4. #04SAM607
  5. #05GPT-4245
  6. #06LLaMA227
  7. #07Stable Diffusion172
  8. #08Whisper147

TITLE FOCUS · BENCHMARKS

  1. #01VQA367
  2. #02ImageNet191
  3. #03MS COCO63
  4. #04PolyGlot38
  5. #05SQuAD31
  6. #06MMLU24
  7. #07SWE-bench20
  8. #08Pass@k16

CATEGORY DISTRIBUTION

CategorySharePapers
cs.CV
156.1K
cs.LG
140.4K
cs.CL
84.6K
cs.AI
41.8K
cs.RO
41.7K
stat.ML
18.5K
eess.IV
15.5K
cs.CR
9.8K
cs.IR
7.6K
cs.SD
6.0K
cs.HC
5.6K
cs.SE
5.3K
cs.CY
5.2K
eess.SP
4.5K

MOST CITED INSIDE THE RESOLVED GRAPH

  1. cs.CL 2023-03-15 19,691 inbound citations
    GPT-4 Technical Report
    PAPER::2303.08774
  2. cs.LG 2014-12-22 12,880 inbound citations
    Adam: A Method for Stochastic Optimization
    PAPER::1412.6980
  3. cs.AI 2024-07-31 12,851 inbound citations
    The Llama 3 Herd of Models
    PAPER::2407.21783
  4. cs.CL 2023-02-27 11,263 inbound citations
    LLaMA: Open and Efficient Foundation Language Models
    PAPER::2302.13971
  5. cs.CL 2023-07-18 11,127 inbound citations
    Llama 2: Open Foundation and Fine-Tuned Chat Models
    PAPER::2307.09288
  6. cs.CV 2020-10-22 9,408 inbound citations
    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
    PAPER::2010.11929
  7. cs.LG 2017-07-20 9,078 inbound citations
    Proximal Policy Optimization Algorithms
    PAPER::1707.06347
  8. cs.CL 2025-01-22 8,546 inbound citations
    DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
    PAPER::2501.12948

NEWEST PAPER OBJECTS

  1. hep-th 2026-07-30
    Learning to Trace Seiberg Dualities
    PAPER::2607.28628
  2. cs.CV 2026-07-30
    ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
    PAPER::2607.28627
  3. cs.CV 2026-07-30
    ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
    PAPER::2607.28625
  4. cs.CV 2026-07-30
    PhiZero: A World Model Built Around Physical Language
    PAPER::2607.28624
  5. cs.RO 2026-07-30
    PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
    PAPER::2607.28623
  6. cs.CL 2026-07-30
    AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
    PAPER::2607.28618
  7. cs.AI 2026-07-30
    AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
    PAPER::2607.28617
  8. cs.CV 2026-07-30
    Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
    PAPER::2607.28611

RESEARCH CODE · TOP GITHUB ORGANISATIONS

  1. #01huggingface4.3K
  2. #02facebookresearch3.3K
  3. #03tatsu-lab2.6K
  4. #04microsoft2.0K
  5. #05openai1.8K
  6. #06meta-llama1.8K
  7. #07google1.6K
  8. #08ultralytics1.5K

CODE GRAPH PROOF

Paper → repository edges245,681
Distinct papers with repository relation136,120
Canonical repositories144,932
Code-positive enriched papers71,213

Repository identity is canonicalised as github.com/<lower-owner>/<lower-repository>. Code-positive classification covers the currently enriched subset and is not presented as full-corpus coverage.

PROVENANCE + COVERAGE

Know what is live.
Know what is incomplete.

Coverage is published as evidence, not hidden behind a generic “AI-powered” claim. Author and institution outputs remain exploration signals where identity resolution is incomplete.

API INGEST606,189

Paper objects in the production database.

HTML PARSED100.0%

606,189 paper HTML records processed.

WHITELIST MATCH100.0%

606,189 paper objects evaluated by the entity layer.

ENRICHED87.4%

529,596 paper objects enriched for empirical and code signals.

QUALITY::WARNING · LAST_QUALITY::2026-08-09T13:45:05Z · FRONTIER::2007-03-07T00:00:00Z · AFFILIATION_RELATIONS::87,819 · NORMALIZED_LABELS::17,949

MCP ACCESS

The graph becomes useful
when your agent can query it.

Search, resolve, compare, trace and connect research objects through the live BrunoSan ArXiv MCP interface.

ACCESS

Start with one domain.
Connect the whole picture.

One transparent price per domain. LENS connects them. Full Intelligence unlocks everything.

TRIALNO CARD

Free Trial

€0

Prove the value with real BrunoSan intelligence.

  • 100 MCP calls
  • One selected intelligence domain
  • Real objects, sources and relationships
  • Valid for 365 days
  • The same stable API key when you upgrade
  • No credit card required
Create free API key
ONE DOMAINFOUNDING PRICE

Domain Access

€19.95

/ month

2 months free on yearly billing

Go deep in the domain that matters now.

  • Choose any current intelligence domain
  • All MCP tools available in that domain
  • 90-day history
  • Unlimited calls at 30 requests per minute
  • Priority support
Start Domain Access
BEST VALUELENS INCLUDED

Full Intelligence

€199

/ month

Save €50.40/month vs. individual access

Your agent sees every relevant domain—not only the one you expected.

  • All 11 intelligence domains
  • LENS cross-domain connection layer
  • All 12 MCP endpoints
  • Cross-domain entity intelligence
  • 90-day history across all domains
  • Unlimited calls at 30 requests per minute
Get Full Intelligence
ENTERPRISECUSTOM

Enterprise

Custom

Deploy BrunoSan around your data, controls and operating model.

  • Private LibertyOS instance
  • Dedicated EU deployment
  • Custom entity tracking
  • RBAC, SLA and DPA
  • White-label or on-premise options
Book an enterprise demo

LENS ADD-ON

Connections are the premium.

Resolve the same entity across your active domains and expose the combined signals, sources and relationships in one traceable view.

€29.95/ monthRequires 2+ active domains
Add LENS
Cancel any timeMonthly billing, no minimum term, no setup fee. Access runs to the end of the paid period.
Founding price protectionFounding customers keep their subscription price for at least 24 months.
Transparent billing2 months free on yearly access. No hidden per-tool surcharge.
Controlled expansionIf fewer than two domains remain active, LENS stays available until the current billing period ends and will not renew.

FAQ

Research intelligence,
defined clearly.

These answers are visible to readers, Google Search and AI answer systems without requiring JavaScript.

What is the BrunoSan ArXiv research graph?

It is a live graph of addressable AI research objects: papers, authors, extracted entities, citation relations, canonical GitHub repositories and affiliation signals. Every result is connected back to source data.

How is this different from searching arXiv.org?

Search returns documents. BrunoSan resolves relations between documents and exposes those relations through MCP tools for trends, citations, co-occurrence, code, people and institutions.

Can an agent find papers that publish code?

Yes. The graph links papers to canonical GitHub repositories and the search tool can also filter the currently enriched set for papers classified as having code.

Are citation counts external estimates?

No. The visible citation count is derived from reference records whose ArXiv identifiers resolve to papers present in the BrunoSan graph.

Are author and institution rankings definitive?

They are exploration signals. Author-name resolution and affiliation extraction can be incomplete, so role and institution results include quality context rather than claiming a universal academic ranking.

How do I connect the graph to an AI agent?

Use the BrunoSan ArXiv MCP endpoint at https://arxiv.mcp.brunosan.de/mcp. Domain Access includes the current public ArXiv MCP tool registry.