NEWEST · AI Agents
- OpenForgeRL: Train Harness-native Agents in Any Environment
- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
- Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry
- Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks
- Regulating autonomous and agentic AI
- Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation
- Explainability Framework for Policy-Aware Autonomous Agents
- SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
- HARP: The Human--AI Research Platform
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
- The Ethics of Autonomous AI Agents for Offensive Security
- Bayesian uncertainty estimation improves clinical decision making in medical AI agents
- A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace
- DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
- Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
- Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation
- Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
- ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
- Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
- BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
- Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models
- Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts
- SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
- The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
- ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems
- Engineering Trustworthy Agentic AI for Critical Systems
- Operational Hallucination and Safety Drift in AI Agents
- LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
- Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data
- Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?