Researchers have made significant progress in developing proactive service agents, which can plan, invoke tools, and modify external states. A unified decision framework, methods, and evaluation for proactive service agents have been proposed, addressing the challenges of incomplete environmental and user signals. The formulation represents timing, content, and delivery within one structured action, making explicit the option value of waiting, the decision value of questions, and feedback-induced state changes. Existing methods have been organized along a decision pipeline, and prescribed, predictive, model-based, and return-optimizing mechanisms have been described as nonexclusive policy-construction components. Reliable proactive service requires calibrated incremental intervention value, verifiable authorization, recoverable execution, and counterfactual evidence.
A novel mechanism, Speculative Macro Commit (SMC), has been introduced for a two-tier agent system, which reduces latency by 10.23% over the Speculative Actions (SA) baseline and 18.59% over sequential execution on the $\tau^2$-Bench Telecom subset. SMC mines recurring multi-action skeletons from training traces and stores them in a macro library used to match against action chains predicted by the drafter at runtime.
Researchers have proposed a framework for personalizing general-purpose LLM/RAG-based AI teaching assistants across academic disciplines and courses. The framework adapts responses using six learner-specific dimensions and student queries are analyzed using Bloom's Taxonomy to estimate cognitive complexity at the interaction level. The framework is evaluated through experiments using NLP metrics and a human study with five participants, showing perceived differences in response style and structure across personalization conditions.
A benchmark, DuplexSpeechBench-IFEval (DSB-IFEval), has been introduced for evaluating implicit instruction-following in real-time spoken interaction. DSB-IFEval comprises 1,038 test cases spanning eight diverse assistant roles and evaluates five conditioning protocols for instruction-following. The evaluation shows architecture-dependent trade-offs, with full duplex models like F-Actor and PersonaPlex being more sensitive to whether conversational behavior is stated explicitly or must be inferred from a persona.
Key Takeaways
- Proactive service agents can plan, invoke tools, and modify external states, but require calibrated incremental intervention value, verifiable authorization, recoverable execution, and counterfactual evidence.
- Speculative Macro Commit (SMC) reduces latency by 10.23% over the Speculative Actions (SA) baseline and 18.59% over sequential execution on the $\tau^2$-Bench Telecom subset.
- Personalizing general-purpose LLM/RAG-based AI teaching assistants across academic disciplines and courses improves response style and structure.
- Implicit instruction-following in real-time spoken interaction is challenging, with architecture-dependent trade-offs between full duplex models like F-Actor and PersonaPlex.
- Large language models can struggle with complex queries requiring fine-grained visual details or external knowledge, but can be improved with necessary tool-evidence path rewards.
- Neonatal respiratory disease diagnosis requires a knowledge-logic-alignment framework that incorporates neonatologist-inspired diagnostic priors into multimodal representations.
- Multimodal culinary reasoning requires a benchmark that probes the knowledge-application gap, showing that near-perfect recognition can conceal an inability to apply cultural knowledge.
- Counterfactual explanations for the shortest path problem require a runtime mechanism that iteratively incorporates constraints until an exact solution is found.
- Contrastive pretraining of small encoders with synthetically generated natural-language descriptions emphasizing code functionality and intent improves code representation learning.
- Prompt engineering for drug toxicity prediction requires a method to analyze prompt engineering, showing that the natural variance in LLMs outweighs any fine-tuning of prompts.
Sources
- Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
- Speculative Macro Commit for Faster Tool-Using Agents
- MasterControl Seventeen Every Time
- Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
- A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant
- DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
- AutoGraphForge: Towards Automated Graph Theory Discovery
- Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
- GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
- Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
- NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis
- CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
- PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing
- GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
- KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
- HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews
- Counterfactual Routing Using Integer Programming with Constraint Generation
- Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study
- Analysis of Prompt Engineering for Drug Toxicity Prediction
- A computable representation of the physical laboratory enables verifiable workflows
- Transfiver: Human-AI Co-Inference through a Shared Editable State
- Rethinking World Models for Safety-Critical Embodied Systems
- Semantic Bayesian World Models
- SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation
- Inferring Affective Consciousness in an Artificial Agent: A Case Study
- More Criticism Does Not Make a Better Review: EquiReview-R
- Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding
- Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting
- Instruction Duplication as an Inference-Time Control Primitive
- InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models
- The Dually Flat Geometry of Planning as Inference
- Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
- DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
- IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations
- Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
- Efficient Test-Time Adaptation through Human-AI Interaction
- Environment Evolution for Terminal Agents
- Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
- Rethinking On-Policy Distillation of Large Language Models II: One Training Example
- Dalek: A Constructive Agent Machine
- Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation
- A Computationally Feasible Framework for Causal Probabilistic Explanation
- A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
- From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
- The Natural Language Interaction Protocol and Standard for AI Agents
- Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing
- Interface-Induced Trajectory Censoring
- FiMI Banking: A Sovereign Model for Indian Retail Banking
- Value-Preserving Architectures for Agentic AI Systems
- Lose the Order, Keep the Hierarchy: Deordering HTN Plans
- Xiaomi-TabLDM: A Tabular Foundation Model Technical Report
- SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
- Artificial Intelligence for Energy Optimization in Data Centers
- The Attention Triangle in Audio-Video Models
- Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection
- Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
- Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence
- Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable
- Spurious Advantage Hidden in GRPO
- FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models
- LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening
- Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations
- CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception
- Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions
- What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation
- STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation
- Bioinfoysis Technical Report
Comments
Please log in to post a comment.