Researchers have made significant progress in various fields, including AI, language models, and multimodal systems. In AI, a new framework for proactive service agents has been proposed, which enables agents to infer service opportunities from incomplete environmental and user signals. In language models, a new benchmark for situated measurement grounding has been introduced, which evaluates the ability of models to provide accurate and reliable measurements in industrial scenes. In multimodal systems, a new approach for counterfactual routing has been proposed, which uses integer programming with constraint generation to find counterfactual explanations for the shortest path problem. These advancements have the potential to improve the performance and reliability of AI systems in various domains.
A new method for efficient test-time adaptation through human-AI interaction has been proposed, which integrates human-agent interaction signals into agent context and weights. This approach has been shown to improve the performance of agents in various tasks, including writing and visual creation. Additionally, a new framework for lifelong learning AI agents has been introduced, which enables agents to transform experience and accumulated knowledge into durable, reusable competence. This framework has been evaluated on a traffic simulation task and has shown promising results.
Researchers have also made progress in the field of bioinformatics, where a new approach for adapting to evolving requirements has been proposed. This approach uses a graph-constrained agentic framework to adapt to changing requirements in retail supply chain operations. The framework has been evaluated on a real-world dataset and has shown promising results. Furthermore, a new method for efficient test-time adaptation through human-AI interaction has been proposed, which integrates human-agent interaction signals into agent context and weights. This approach has been shown to improve the performance of agents in various tasks, including writing and visual creation.
Key Takeaways
- Researchers have proposed a new framework for proactive service agents that enables agents to infer service opportunities from incomplete environmental and user signals.
- A new benchmark for situated measurement grounding has been introduced, which evaluates the ability of models to provide accurate and reliable measurements in industrial scenes.
- A new approach for counterfactual routing has been proposed, which uses integer programming with constraint generation to find counterfactual explanations for the shortest path problem.
- A new method for efficient test-time adaptation through human-AI interaction has been proposed, which integrates human-agent interaction signals into agent context and weights.
- A new framework for lifelong learning AI agents has been introduced, which enables agents to transform experience and accumulated knowledge into durable, reusable competence.
- Researchers have proposed a new approach for adapting to evolving requirements in retail supply chain operations, using a graph-constrained agentic framework.
- A new method for efficient test-time adaptation through human-AI interaction has been proposed, which integrates human-agent interaction signals into agent context and weights.
- A new framework for lifelong learning AI agents has been introduced, which enables agents to transform experience and accumulated knowledge into durable, reusable competence.
- Researchers have proposed a new approach for adapting to evolving requirements in retail supply chain operations, using a graph-constrained agentic framework.
- A new method for efficient test-time adaptation through human-AI interaction has been proposed, which integrates human-agent interaction signals into agent context and weights.
Sources
- DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
- Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
- PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing
- GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving
- Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
- AutoGraphForge: Towards Automated Graph Theory Discovery
- KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
- HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
- Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection
- A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant
- A Computationally Feasible Framework for Causal Probabilistic Explanation
- A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms
- From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
- The Natural Language Interaction Protocol and Standard for AI Agents
- Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing
- Interface-Induced Trajectory Censoring
- FiMI Banking: A Sovereign Model for Indian Retail Banking
- Value-Preserving Architectures for Agentic AI Systems
- Lose the Order, Keep the Hierarchy: Deordering HTN Plans
- Xiaomi-TabLDM: A Tabular Foundation Model Technical Report
- Dalek: A Constructive Agent Machine
- Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable
- IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations
- Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI
- DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions
- Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation
- NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis
- CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
- What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation
- GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
- A computable representation of the physical laboratory enables verifiable workflows
- The Attention Triangle in Audio-Video Models
- Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation
- Artificial Intelligence for Energy Optimization in Data Centers
- Counterfactual Routing Using Integer Programming with Constraint Generation
- Transfiver: Human-AI Co-Inference through a Shared Editable State
- Rethinking World Models for Safety-Critical Embodied Systems
- More Criticism Does Not Make a Better Review: EquiReview-R
- Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding
- Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting
- Instruction Duplication as an Inference-Time Control Primitive
- InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models
- The Dually Flat Geometry of Planning as Inference
- FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models
- Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
- DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
- Spurious Advantage Hidden in GRPO
- Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
- Efficient Test-Time Adaptation through Human-AI Interaction
- Environment Evolution for Terminal Agents
- Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints
- Rethinking On-Policy Distillation of Large Language Models II: One Training Example
- SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation
- Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study
- Analysis of Prompt Engineering for Drug Toxicity Prediction
- LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening
- STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation
- Bioinfoysis Technical Report
- Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations
- CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception
- Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
- Speculative Macro Commit for Faster Tool-Using Agents
- MasterControl Seventeen Every Time
- Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence
- Semantic Bayesian World Models
- SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation
- Inferring Affective Consciousness in an Artificial Agent: A Case Study
Comments
Please log in to post a comment.