Frontier models show divergent error rates despite similar success; OpenDiscoveryTrace reveals Claude Opus 4.6 generates 30× more errors than GPT-5.4. Decision-Focused Active Learning optimizes NdFeB magnet recovery in 16–24 experiments versus 48 for nonadaptive methods. State-Path Tool Menus boost ToolBench success from 0.737 to 0.898, while RobustSGPO raises held-out task completion from 60% to 80%. TrimSFT improves mathematical reasoning by +26.9 points on MATH500 via logit-gap reweighting. Autonomous GeoAI agents integrate ecological criteria, and Multi-Agent Agentic Graph Learning partitions graphs better than SOTA.
Safety and reliability research exposes critical gaps: a black-box audit found 85% behavior vulnerabilities in CrewAI and AutoGen. Trust evaluation frameworks show top models score 95.1 on task completion but only 22.6 on safety. LexAgentHallu profiles cascading legal hallucinations invisible to outcome metrics, while RESCUE-BENCH demonstrates LLM failures in relation-sensitive multi-party support. CareGuard detects cyberbullying efficiently, and OntologyAligner achieves 88.78% accuracy in biomedical normalization. Proof-carrying cognition proposals aim to close verification gaps, though unsound verifiers show soundness loss.
Long-horizon tasks benefit from subagents with clear contracts, outperforming context-loading skills despite overhead. Adaptive entangled game modules explain 82–94% of trading patterns, supporting nonlocal brain hypotheses. Cyber-financial contagion models reveal AI vendor compromises propagate through banking networks with AUROC 0.82 early warnings. Memory management improves via storage-usage separation, boosting personalization scores by 2.4–4.0 points and reducing latency by 15–61%. CityPlanner uses sandbox agents and atomic-task RL to solve urban planning, surpassing baselines.
New architectures address agent limitations: a four-layer Cognitive Digital Twin enables self-evolving operations, and a Belief-State Engine augments LLMs with Bayesian posteriors for partial observability. ConvMem uses hierarchical convolution for efficient long-context reasoning without RL overfitting. Fortunate Recall applies behavioral ontologies to manage memory lifecycles, reducing confabulation by half. SmartWeatherAgent fuses rules with LLMs for weather alerts, boosting warning quality by 112%. TRACE uses synthesized rewards to improve diagnostic reasoning, and Valerant generates navigable 3D game maps using action-conditioned world models.
Key Takeaways
- Claude Opus 4.6 produces 30× more errors than GPT-5.4 despite similar success rates.
- Decision-Focused Active Learning reduces NdFeB magnet recovery experiments to 16–24.
- State-Path Tool Menus increase ToolBench success from 0.737 to 0.898.
- RobustSGPO improves agent harness evolution completion from 60% to 80%.
- TrimSFT gains +26.9 points on MATH500 via logit-gap reweighting.
- Black-box audits find 85% behavior vulnerabilities in CrewAI and AutoGen.
- Trust frameworks show 95.1 task completion vs 22.6 safety scores.
- Subagents outperform context-loading skills with clear input-output contracts.
- Memory separation boosts personalization by 2.4–4.0 points and cuts latency 15–61%.
- Fortunate Recall reduces memory confabulation by half via behavioral ontologies.
Sources
- OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
- Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery
- The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents
- An Autonomous GeoAI Agent for Arctic Eco-Navigation
- Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
- XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?
- Multi-Agent Agentic Graph Learning via Structural Signatures
- A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
- RobustSGPO: Search-Space Control for Agent Harness Evolution
- Seven Sources of Physical AI Capability Formation
- Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
- Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation
- Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
- PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations
- RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
- UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model
- Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
- Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks
- LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
- Structural Process Supervision for Latent Chain-of-Thought Reasoning
- Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models
- Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications
- Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format
- Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts
- RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases
- Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
- From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
- TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
- Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection
- Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition
- ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance
- Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
- Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions
- Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks
- Adaptive Entangled Game Modules in Artificial General Intelligence
- Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System
- What Should an Agent Forget? Separating What Is Stored from What Is Used
- Kernel-Managed Shared Memory for System-Wide Personalization
- AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
- Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields
- The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents
- Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety
- OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization
- Grounded Evaluation and Repair for NL-to-PDDL Problem Generation
- CityPlanner: A Sandbox Agent for Executable Urban Planning
- Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
- From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins
- Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability
- ConvMem: Convolutional Memory for Long-Context Reasoning
- Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs
Comments
Please log in to post a comment.