VERTICAL::ARXIV · LIVE

OBJECT::RESEARCH_TRACK · ai-agents

AI Agents.
Chronological research intelligence.

Research on autonomous and agentic AI systems, tool-using agents and agent-oriented evaluation, excluding papers already classified as multi-agent systems.

4.9KTrack papersAll deterministic members
1.3KTitle focusDefining entity appears in title
772With codeEnriched papers with code signal
2007-09-15Oldest paperEarliest member in the graph
2026-07-23Newest paperLatest member in the graph
8Defining entitiesStable entity objects in the track definition

TRACK_ID::f5ec2268-4c6e-57d5-839c-d58e4abf6065 · JSON::https://brunosan.de/arxiv/tracks/ai-agents.json · MCP::https://arxiv.mcp.brunosan.de/mcp

CHRONOLOGICAL VIEW

Newest research.
Stable paper objects.

Sorted by publication date and ArXiv ID. Every result remains the same PAPER object used everywhere else in the graph.

NEWEST · AI Agents

  1. cs.AI 2026-07-23 reinforcement learning
    OpenForgeRL: Train Harness-native Agents in Any Environment
    PAPER::2607.21557
  2. cs.AI 2026-07-23 reasoning
    Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
    PAPER::2607.21503
  3. cs.AI 2026-07-23 other
    Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry
    PAPER::2607.21495
  4. cs.AI 2026-07-23 code generation
    Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks
    PAPER::2607.21482
  5. cs.AI 2026-07-23 other
    Regulating autonomous and agentic AI
    PAPER::2607.21345
  6. cs.CR 2026-07-23 reasoning
    Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation
    PAPER::2607.21325
  7. cs.LO 2026-07-23 planning
    Explainability Framework for Policy-Aware Autonomous Agents
    PAPER::2607.21209
  8. cs.AI 2026-07-23 reasoning
    SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
    PAPER::2607.20926
  9. cs.HC 2026-07-22 other
    HARP: The Human--AI Research Platform
    PAPER::2607.20773
  10. cs.AI 2026-07-22 code generation
    NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
    PAPER::2607.20709
  11. cs.CR 2026-07-22 other
    The Ethics of Autonomous AI Agents for Offensive Security
    PAPER::2607.20255
  12. cs.LG 2026-07-22 image classification
    Bayesian uncertainty estimation improves clinical decision making in medical AI agents
    PAPER::2607.20582
  13. cs.HC 2026-07-22 other
    A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace
    PAPER::2607.19941
  14. cs.AI 2026-07-22 reasoning
    DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
    PAPER::2607.19865
  15. cs.AI 2026-07-22 other
    Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
    PAPER::2607.19837
  16. cs.AI 2026-07-22 reasoning
    Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation
    PAPER::2607.19793
  17. cs.CL 2026-07-22 other
    Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
    PAPER::2607.19747
  18. cs.AI 2026-07-21 other
    ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
    PAPER::2607.19321
  19. cs.AI 2026-07-21 other
    Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
    PAPER::2607.19297
  20. cs.AI 2026-07-21 reasoning
    BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
    PAPER::2607.19262
  21. cs.LG 2026-07-21 other
    Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models
    PAPER::2607.19453
  22. cs.SE 2026-07-21 other
    Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts
    PAPER::2607.18970
  23. cs.AI 2026-07-21 question answering
    SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
    PAPER::2607.18665
  24. cs.AI 2026-07-20 reasoning
    The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
    PAPER::2607.19433
  25. cs.CR 2026-07-20 other
    ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems
    PAPER::2607.19432
  26. cs.AI 2026-07-20 survey
    Engineering Trustworthy Agentic AI for Critical Systems
    PAPER::2607.18548
  27. cs.AI 2026-07-20 planning
    Operational Hallucination and Safety Drift in AI Agents
    PAPER::2607.18366
  28. eess.SY 2026-07-20 reasoning
    LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
    PAPER::2607.18147
  29. cs.SE 2026-07-20 other
    Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data
    PAPER::2607.18064
  30. cs.CR 2026-07-20 other
    Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
    PAPER::2607.17986

MOST CITED · AI Agents

  1. cs.CL 2022-10-06 2,383 inbound citations
    ReAct: Synergizing Reasoning and Acting in Language Models
    PAPER::2210.03629
  2. cs.AI 2023-07-25 872 inbound citations
    WebArena: A Realistic Web Environment for Building Autonomous Agents
    PAPER::2307.13854
  3. cs.AI 2023-07-31 842 inbound citations
    ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
    PAPER::2307.16789
  4. cs.CV 2023-10-03 791 inbound citations
    MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
    PAPER::2310.02255
  5. cs.AI 2023-08-07 752 inbound citations
    AgentBench: Evaluating LLMs as Agents
    PAPER::2308.03688
  6. cs.AI 2024-08-12 711 inbound citations
    The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
    PAPER::2408.06292
  7. cs.CL 2025-04-28 634 inbound citations
    Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
    PAPER::2504.19413
  8. cs.AI 2024-06-17 587 inbound citations
    $τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
    PAPER::2406.12045

TRAJECTORY

How the field
moves over time.

Volume is counted from track membership, not keyword frequency. The side panels show what the track connects to inside the same graph.

PAPERS BY YEAR

YearVolumePapers
2007
1
2011
4
2012
3
2013
8
2014
6
2015
8
2016
18
2017
40
2018
77
2019
97
2020
93
2021
149
2022
149
2023
289
2024
610
2025
1,574
2026
1,822

TOP CONNECTED ENTITIES

  1. #01RAG183
  2. #02Claude179
  3. #03VLM124
  4. #04Chain-of-Thought121
  5. #05GPT-4o110
  6. #06Gemini108
  7. #07LLaMA96
  8. #08SFT95
  9. #09Zero-Shot94
  10. #10GPT-489
  11. #11Interpretability87
  12. #12Alignment81

RESEARCH CODE · TOP GITHUB ORGANISATIONS

  1. #01langchain-ai121
  2. #02significant-gravitas86
  3. #03microsoft81
  4. #04openclaw74
  5. #05openai73
  6. #06huggingface66
  7. #07yoheinakajima36
  8. #08meta-llama35
  9. #09google31
  10. #10anthropics31

TRACK DEFINITION

Object IDf5ec2268-4c6e-57d5-839c-d58e4abf6065
Slugai-agents
Kindresearch_theme
Match modeentity
Title focus1,309

in_title=1 means at least one defining entity is present in the paper title.

ACCESS

Start with one domain.
Connect the whole picture.

One transparent price per domain. LENS connects them. Full Intelligence unlocks everything.

TRIALNO CARD

Free Trial

€0

Prove the value with real BrunoSan intelligence.

  • 100 MCP calls
  • One selected intelligence domain
  • Real objects, sources and relationships
  • Valid for 365 days
  • The same stable API key when you upgrade
  • No credit card required
Create free API key
ONE DOMAINFOUNDING PRICE

Domain Access

€19.95

/ month

2 months free on yearly billing

Go deep in the domain that matters now.

  • Choose any current intelligence domain
  • All MCP tools available in that domain
  • 90-day history
  • Unlimited calls at 30 requests per minute
  • Priority support
Start Domain Access
BEST VALUELENS INCLUDED

Full Intelligence

€199

/ month

Save €50.40/month vs. individual access

Your agent sees every relevant domain—not only the one you expected.

  • All 11 intelligence domains
  • LENS cross-domain connection layer
  • All 12 MCP endpoints
  • Cross-domain entity intelligence
  • 90-day history across all domains
  • Unlimited calls at 30 requests per minute
Get Full Intelligence
ENTERPRISECUSTOM

Enterprise

Custom

Deploy BrunoSan around your data, controls and operating model.

  • Private LibertyOS instance
  • Dedicated EU deployment
  • Custom entity tracking
  • RBAC, SLA and DPA
  • White-label or on-premise options
Book an enterprise demo

LENS ADD-ON

Connections are the premium.

Resolve the same entity across your active domains and expose the combined signals, sources and relationships in one traceable view.

€29.95/ monthRequires 2+ active domains
Add LENS
Cancel any timeMonthly billing, no minimum term, no setup fee. Access runs to the end of the paid period.
Founding price protectionFounding customers keep their subscription price for at least 24 months.
Transparent billing2 months free on yearly access. No hidden per-tool surcharge.
Controlled expansionIf fewer than two domains remain active, LENS stays available until the current billing period ends and will not renew.

FAQ

Track semantics,
defined clearly.

The definition, object ID and counts are server-rendered from the production graph and exposed as JSON as well.

What is the AI Agents research track?

It is a stable BrunoSan research-track object built from existing paper-to-entity relations. Research on autonomous and agentic AI systems, tool-using agents and agent-oriented evaluation, excluding papers already classified as multi-agent systems.

How are papers assigned to this track?

Membership is deterministic and derived from already resolved entity objects. The track does not run a second crawler, LLM classifier or hidden semantic search.

What does title focus mean?

Title focus counts papers where at least one defining track entity appears in the paper title. It is the stricter relevance subset; the chronological list contains all track members.

Can an AI agent query the same research view?

Yes. The BrunoSan ArXiv MCP endpoint is https://arxiv.mcp.brunosan.de/mcp; track-specific MCP tools use the same stable track slug and object ID.