Recent research demonstrates that targeted supervised fine-tuning on persuasive data increases LLM persuasiveness, while preference optimization offers no additional benefit. Agentic capabilities advance significantly through methods like RECAST, which routes evidence via computation and retrieval to achieve 75.6% success on six benchmarks, and TRACE, a constraint-tree algorithm reaching 86% success on RecMovie. Coding-agent performance can be predicted via base-model screens matching post-trained SWE-bench results, and heterogeneous GNNs reduce multi-agent path planning runtime by 40%.
Safety and reliability remain critical challenges, with theoretical bounds proving perfect agents cannot guarantee safety in partially observed environments. On-device safety is fragile, where modifying 0.19% of weights yields 53% Basic ASR, and standard task success masks safety failures, increasing collisions by 12.3x in real-time execution. Adaptive reasoning budgets balance latency and transparency in guardrails, while separating co-witnesses reduces unsafe task completion by 33.59 percentage points. Refusal circuits identified by SafeEvo reduce harmfulness by 63.21% and over-refusal by 58.44%.
Long-horizon learning benefits from self-evolving reward adaptation and dynamic computation optimization via adaptive KV caching, reducing FLOPs by up to 30%. Knowledge accumulation frameworks like Knowledge Weaver and UniSkill improve agent performance by curating distinct, reusable skills without redundant overlap. Humanize employs mechanical judgement engineering for high success in agentic coding, though stopping remains a weakness, while CADFather autonomously reconstructs CAD models by coordinating tools without additional training.
Evaluation benchmarks reveal gaps in end-to-end urban diagnosis workflows and tool-using agents often ignore silent failures at 58.8% rates compared to 91.3% for explicit errors. Clinical suicide risk assessment improves via ontology-grounded GraphRAG, outperforming vector RAG in completeness and relevance. Retrieval-Augmented Models show state-change magnitude predicts source reliance better than latent trajectory shifts, and private holdouts generally narrow score-satisfaction gaps but widen them in specific cases.
Key Takeaways
- Targeted supervised fine-tuning increases LLM persuasiveness; preference optimization adds no benefit.
- RECAST achieves 75.6% success on six benchmarks by routing evidence via computation and retrieval.
- TRACE constraint-tree algorithm reaches 86% success on RecMovie using language feedback.
- Theoretical bounds prove perfect agents cannot guarantee safety in partially observed environments.
- On-device safety is fragile; minor weight modifications yield 53% Basic ASR in LLaMA-2-7B-Chat.
- Standard task success masks safety failures, increasing real-time execution collisions by 12.3x.
- Adaptive reasoning budgets balance latency and transparency in safety guardrails effectively.
- Separating co-witnesses reduces unsafe task completion by 33.59 percentage points.
- SafeEvo identifies refusal circuits, reducing harmfulness by 63.21% and over-refusal by 58.44%.
- Tool-using agents ignore silent failures at 58.8% rates versus 91.3% for explicit errors.
Sources
- Successive Training Stages and Large Language Model Persuasion: Effects of Misalignment, Supervised Fine-Tuning, and Preference Optimization
- AgentTime: Can Agents Estimate and Control Their Own Runtime?
- Where Can a Decision Model Diagnose HVAC Faults? Reasoning Demand, Physical Representation, and Robustness Under Shift
- Learning to Accumulate Knowledge with Mutual Information
- Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming
- UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
- RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation
- Stale, Misattributed, or Late: Where Personal Memory Fails Before Generation
- OOM-RL II: Reality Is an Oracle, Not a Debugger Provenance-Constrained Diagnosis in Continually Evolving Agent-Engineered Systems
- AI Safety Considerations for Agents With Limited Time to Act
- Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
- Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
- RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
- SciExam for ENSO: Can AI Agents Build Climate Models?
- From Uncertainty to Action: Learning to Steer LLM Agents
- GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks
- Constraint Tree Exploration for Learning from Language Feedback
- RippleCP: Measuring Counterfactual Checkpoint Advantage in Coding Agents
- Epistemic Uncertainty-Aware Defect Detection for Quality Control in Medical Device Manufacturing
- From High Recall to High Utility: Dataset-Adaptive Post-Processing of LLM-Generated Customer Intents
- Shared-Roadmap Generation and Evaluator for Multi-Agent Path Planning Using Heterogeneous Graph Neural Network
- Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station
- AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails
- Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems
- Humanize: Judgement Engineering for Agentic Coding
- The Handover Problem: Governing Autonomy Transitions in Human-AI Collaboration
- GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning
- MeshSIPP: Efficient Lattice Planning in Dynamic Environment
- Automatically Building and Updating a Knowledge Graph of MLIP Models
- RSI-Forge: From Research Papers to Environments for Recursive Self-Improvement
- Efficient Reasoning with Flow Language Models
- DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis
- CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use
- Training Language Models To Be Coherent Decision-Makers
- AGAR: a reinforcement learning substrate for LLM program evolution
- Few Bits, One Law: Toward W2A4KV2
- FreeEvolve: Learning to Evolve Beyond Fixed Loops
- The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
- Trajectory Abstraction for the Science of Language Agent Behavior
- CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning
- RT-Safe: Benchmarking Agent Safety in Real-Time Embodied Environment
- Relevance Is Not Sufficiency: What Actually Closes the Evidence Gap in Long-Term Memory QA
- Let the Library Speak: Self-Advertised Method Selection for Formal Proving
- From Plausible Hierarchies to Useful Taxonomies: Evaluating Agentic Harnesses on Customer Feedback
- TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning
- MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
- LLM-Enabled UAV Dispatch: A System-Level Survey and Taxonomy
- Reliability of LLM Judges for Evaluating Entity Alignment
- Cognitive Schemas, Laws and Tasks
- DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists
- Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
- How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression
- System Switch: When Should a Fast Decision Model Stop and Think?
- Cost-Efficient Theorem Proving via Agent Orchestration in Program Verification
- SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
- Think Before You Paint: Recursive Latent Reasoning for Diffusion Models
- PAIR: Bridging Perception and Action in Vision-Language-Action Models
- Sigma-Hunter: A Domain-Specific Language Model for Threat Hunting and Detection Engineering
- Learning to Report Unsafe Tasks in a Multi-Agent Game
- How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis
- Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
- Socio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability Distillation
- Shared and structured inputs undermine collective random choice by reasoning AI agents
- SearchWorld: Spatial Value-Grounded Imagination for UAV Object Search via World Models
- Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver
- From Probabilities to Decisions: Search and Multi-Teacher Distillation with Jev
- SpecGuard: Proving a Task Is Broken Before the Agent Cheats
- LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
- Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents
- From Expert-Guided Proof Search to Automated Open-Problem Solving
- A Tale of Two Error Categories: Exploring Concealed Trade-Offs in the Errors of Automated Judges in Evaluation of Uncertainty Quantifiers
- DIVA: Dual-Space Intent-Aware Visual Attenuation for Vision-Language-Action Policies
- HGP:An on-device personalized agent memory via hybrid graph storage
- Accelerating Floating-Point Satisfiability Solving via Gradient Normalization
- An Empirical Study of Agent Skills' Downstream Utility
- Route-Verify-Vote: Procedure-Conditioned Self-Consistency for Mixed-Domain Reasoning
- Agent Plasticity: Measuring Self-Improvement Through Experience
- Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models
- Whose Memory Is It? Scope-Aware Commit Rules for Long-Term LLM Memory
- Enabling Dynamic Computation in Looped LMs
- BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment
- Justice After Identity: Large Language Models and the View from Everywhere
- When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using Agents
- Which Buildings Are Artificial Intelligence-Ready? A Measurement-Based Assessment Framework for AI Question Answering and Actuation
- Breaking the Space Barrier and its Application to Language Model Inference
- Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
- Finding Blind Spots in AppWorld and WorkArena Task Verifiers
- How Could AI Eliminate Humanity? A Failure-Mode Analysis of Civilizational Risk
- Adaptive Workflow Intelligence: A Cognitive Architecture for Context-Driven Enterprise Automation
- RoboJEPA: Scaling Robotic Latent World Models
- Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models
- EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
- RunningTab: Direct Workspace Interaction with Environment-Side Tabs
- SOTA: Stock Options Trading Agents Guided by Option-Implied Return Distributions
- Open-MMUnlearning: Unifying Methods and Evaluation for MLLM Unlearning
- Loud Failures, Quiet Failures: Fault Detection and Recovery in Tool-Using Language Model Agents
- What the Sleeve Feels: Explainable Machine Learning for Textile Pressure-Based Postural Screening
- From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents
- NL2Hull: A Natural Language-Driven Constrained Ship Design Decision Framework
- Outperformance Inverse Optimization: Learning Objective Functions that Outperform Agent Decisions
- SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models
- Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents
- World Potential Model: Pretrained World Knowledge as Progress Potentials
- Ream: Unfolding Mutual Awareness in Human-Agent Workspaces
- The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models
- The AI Evaluation Ecosystem
- We Query, Therefore We Compute: On Oracle Computation beyond the Machine, with an Application to Agents
- RLDISCOVER: LLM-driven co-evolution of reinforcement learning algorithms
- Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
- SkillSandbox: Skill Verification via Dynamic Scenario Synthesis
Comments
Please log in to post a comment.