Recent advances in LLM efficiency leverage LoRA fine-tuning, Fisher-information pruning, and ternary models to enhance performance while addressing agentic system risks like decision laundering and data leakage. Safety efforts now emphasize trace-grounded evaluation and segment-aware alignment, though metrics often obscure high-severity failures in domains ranging from clinical trials to algorithmic trading. Governance challenges persist, requiring auditable safety cases and legal alignment before policy deployment.
AI safety gaps remain critical, with Safe RL demanding multi-metric reporting and agent-tool anomalies stemming from missing transactional semantics. Medical unlearning requires class-adaptive scoring, while virtual memory masks sensitive spans more effectively than whole-field replacement. Synthetic finetuning fails against reward hacking, and VLM selection must balance cost against accuracy to mitigate visual QA explanatory deficits.
Multimodal and architectural innovations include 3D Gaussian Splatting for talking heads, audio-visual navigation via Transformers, and world model analysis of LLM trajectories. Frameworks like 'ground-truth-as-code' and 'Blindspot' benchmarking support verification, while KV reuse reduces prefill costs and conformal prediction aids regression. Applications span genomic datasets reaching 96.5% accuracy, evacuation routing, and industrial control systems.
Key Takeaways
- LoRA fine-tuning and Fisher-information pruning boost LLM efficiency.
- Agentic systems face risks of decision laundering and data leakage.
- Trace-grounded evaluation reveals metrics often hide high-severity failures.
- Safe RL requires multi-metric reporting to ensure robustness.
- Agent-tool anomalies stem from missing transactional semantics.
- Virtual memory masks sensitive spans better than whole-field replacement.
- Synthetic finetuning fails against reward hacking attacks.
- 3D Gaussian Splatting enables advanced talking head generation.
- KV reuse significantly reduces LLM prefill computational costs.
- Genomic datasets now achieve 96.5% accuracy in analysis.
Sources
- LoRA Fine-Tuned Models for Control Systems Course Q\&A: A Multidimensional Evaluation of Model Scale and Rank Effects
- MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
- Artificial intelligence and biosecurity: capabilities, threat pathways, and defense-in-depth governance
- Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving
- Position: AI Is Not Ready for Strategic Conflicts
- Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation
- Optimal Pruning for Neural Architectures using Fisher Information Distances
- The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
- Metacognitive Steering: Learning the Structure of Scientific Judgment
- Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions
- CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design
- Cross-Anatomy Transfer Versus Sparse Interpolation in Digital-Twin-Oriented Aortic Fluid-Structure Interaction Surrogates
- CLEAR: Cross-Source Evidence Adjudication for Large Language Models in Medicine
- The AI-Enabled Scientific Frontier
- Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs
- Breaking the 1.58-bit Barrier for Ternary LLMs
- Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models
- Query-Aware Source-Risk Triage for Retrieval-Augmented Generation
- QueryFormer: Winning Solution for KDD Cup 2026 Tencent UniRec Challenge
- AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
- Beyond Episodic AI: Cognitive Field Networks for Biologically Inspired Persistent Cognition
- LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture
- VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs
- Integrating the Analytic Hierarchy Process with Large Language Models for Transparent Multi-Criteria Decision-Making
- Turn-level Multiscale Density Ratio Estimation for LLM Agents
- When Should a World Model Move? Loss-Conditioned State Execution
- Navigating Sparse Evidence: Agentic Visual RAG via Explicit Context Selection and Consolidation
- KnowBench: Effort Reduction as a Unified, Deployment-Grounded Benchmark for Clinical AI
- EvoOntology: A Self-Evolving Ontology Layer for Data Agents
- Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science
- Pilot Early, Commit Late: A Real-Options Model of Enterprise AI Adoption under Rapid Technological Progress
- LongAgent: History-Guided Agentic Search for Longitudinal Outcome Prediction
- AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery
- Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
- Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election
- VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries
- STHMoE: Hypergraph-Enhanced Heterogeneous Dependency Coordination for LLM-Based Urban Traffic Data Forecasting
- Medical Knowledge Simplification for Patients in the Era of LLMs: A Case Study on Diabetes
- ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models
- Towards a knowledge-enhanced single-cell foundation model
- Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference
- Recurrent GraphNeural NetworkswithSet-BasedAggregation
- Atria Dawn: The Dawn of Agentic Superintelligence
- Diversified and Perceptible Counterfactual Examples Leveraging Expert Knowledge
- Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA
- Beyond Safe Answers: Segment-Aware Listwise Alignment for Reasoning Safety in Large Reasoning Models
- HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments
- Can AI systems have free will?
- Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning
- CWM: Controllable White-Box Meta-Prompting for Adaptive Retrieval-Augmented Generation and Reasoning Ability
- BusMA: A Bus Communication Substrate for Multi-Agent Systems
- The average-farmer illusion in language-model simulations of agricultural decisions
- AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing
- Moral Rebel Agents: Decision-Making Under Conflicting Obligations
- Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
- DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents
- AI Deployment Accountability Engineering: A Vision for Accountable AI in Safety-Critical Socio-Technical Systems
- Domain Generalization for Smartphone-Based Human Activity Recognition: A Systematic Analysis of Components and Interactions
- El Agente Potente: High-Throughput Agentic Atomistic Simulations
- Token Efficient Task Execution via Application Behavior Modeling for Web Agents
- ANASSA: An Agentic AI Orchestration Framework for Spatial Intelligence
- Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accounting Tasks
- Semantic Knowledge Technologies: what the Semantic Web lost sight of, and what it never had
- VeriDx: Earning the Right to Diagnose with Disease-Centric Verification
- Map Users and Mapmakers: The Scope of Cognitive Attribution from Acquired Representations
- UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics
- IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives
- Drift-Constrained Optimization: Only Direction Matters in Fine-Tuning Instruct Models
- Windowed A-K-MDP
- Enhancing Event Candidate Acquisition for Event Linking
- Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems
- Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic Management
- OrchSLM: Probing the Dynamics of Small Language Model Orchestration
- Design of a Deep Learning Credit Risk Early Warning System Integrating Multi-source Heterogeneous Data
- RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
- MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving
- Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
- Data storytelling meets interpretable machine learning: Decoding AI decisions for non-experts without revealing sensitive data and model details
- New Conditions for Philosophers to Catch the Wave of Citizen Deliberation in the Age of Artificial Intelligence in advance
- When does a scaling result justify a different allocation? A critical review of resource-allocation evidence for AI systems
- FedV-KGQA in Practice: Design Lessons and an Interactive Prototype
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
- ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
- Converge Then Diversify: Decoupling Convergence and Diversity in Multi-Objective Bayesian Optimisation
- LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents
- Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?
- Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents
- From Ideas to Actions: A Public-Data Decision-Support Toolchain Across the Venture Lifecycle
- Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports
- Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement
- Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
- TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models
- AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web Agents
- Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
- From Legal Text to AI-specific Risk Sources: A Systematic Analysis of the EU AI Act's High-Risk Requirements
- Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
- FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks
- How User-AI Mistreatment Occurs and Matters in Conversational Systems?
- Causal multi-modal AI for personalized chemosensitivity prediction
- A Hybrid Agentic AI Framework for Intelligent Supply Chain Analytics
- Cost Characterization of Vertically Partitioned Federated Knowledge Graphs
- Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself
- Solar Intelligence
- Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents
- $\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents
- Recoverability as a System Primitive for Long-Horizon AI Agents
- GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents
- MANAS-2: Constrained Reconstruction for EEG Foundation Models
- JaxAHT: A JAX-Based Library for Ad Hoc Teamwork
- Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks
- Positioning manuscripts in the scientific landscape with agentic AI
- How Many Thoughts Can a Vector Hold? The Capacity of Reasoning by Superposition
- Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges
- Do Not Restart: Residual Completion for Stateful Agent Handoffs
- Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection
- Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs
- ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information
- ViperQ: Order Flow Pattern Recognition via Auction Market Theory for Reinforcement Learning Trading
- Schizophrenia Detection from EEG Signals: A Transformer Framework with Spectrogram Representation
- Convergent Emergence of In-Context Learning Across Modalities
- Synthetic Data in Marketing Research: How to Evaluate and When to Trust
- SAILOR: Solver-Assisted Interactive LLM-based Optimization Recovery
- Dynamic Learning Solutions: A System for Personalized Educational Video Generation
- MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents
- Safety Signals to Verify NetOps Agents with Action-Level Granularity
- Question's Gambit: The First Move Matters in Agentic Deep Search
- A note on goal-based hierarchical RL
- OptoAgent: A Trustworthy Multi-Agent Framework for Opportunistic Vision Micro-Screening in Classroom Environments
- Beyond Scene Description: Multi-Agent Orchestration for Non-visual Access to Virtual Worlds
- Depth and Scale in the Sub-150M Regime: JugnuLM-53M vs JugnuLM-110M
- Diffusion-Based Generation of Gait Trajectories
- Bayesian Intelligence from the Outside
- Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?
- AI Persuasion as a Threat to Human Control
- Converting Sequenced Fuzzy Cognitive Maps to Causal Virtual Worlds with Large Video Generators
- One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling
- Geometric Flow enhanced Graph Coarsening
- GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems
- Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair
- Horizon-specific Expert Fusion for Photovoltaic Power Forecasting
- Four Ledgers, Not One Score: Responsible Communication of LLM-Judge Calibration in Biomedical ML
- Overflip: Repetition-Induced Label Flips in Guardrail Models
- CoMem: Collective-Individual Memory Synergy for Evolutionary Multi-Agent Systems
- OpenAI4S: Code as Action, Science as Sessions
- Enabling Creative Exploration for Vibe Design Agents
- T-LoopFormer: Token-Level Elastic-Depth Looped Transformers for Latent Reasoning With Dynamic Routing
- HazardAuditor: From Executable Threats to Safer Computer-Use Agents
- Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
- Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generation in ESG Domain
- ProIQA: A Process-Based Framework for Fine-Grained Math Item Quality Assessment
- Parameter-Efficient Adaptation of Pretrained Language Models for Time-Series Forecasting
- Evaluation Metrics for Safe Reinforcement Learning
- When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary
- SkillLift: Learning Dense Rubrics from Sparse Oracles for Efficient Skill Evolution
- The Troy Moment of AI: Why Some Will Cheat and Some Will Follow?
- Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents
- GRIN+: Towards Fast Yet Effective Machine Unlearning for Imbalanced Medical Data
- Potential of Artificial Intelligence Algorithms for Identification of Relevant Diagnostic and Prognostic Biomarkers of Early-Stage Liver Cancer
- Semantic-TVM: Structure-Preserving Trustworthy Virtual Memory for Memory-Augmented and Tool-Using Agents
- Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking
- Beyond Accuracy: Robustness, Cost, and Governance Trade-offs for Vision-Language Models in Templated Document Extraction
- NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities
- EEG-Xplain: Decoding Neural Black-Boxes of EEG Foundation Models
- AcquireBound: Runtime Authorization for Resources Acquired by AI Agents
- Are LLMs Good Financial User Simulators? A Preliminary Study
- Predicting build orientation for SLM dental parts: a comparison of rotation representations and direct vector regression
- Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture
- LLM-Enhanced Multi-Agent Reinforcement Learning for Unified Electric Vehicles-Charging Station-Grid Optimization in Public Charging Systems
- Homeostatic Continual Learning
- Planning or Learning: Reliability and Cost in Multi-Asset Maintenance
- QART: A Quantum-Classical Hybrid Architecture for Long-Horizon Reasoning -- Exploring a Conditional Path toward Quantum Scaling
- CoAdapt: An LLM-based Framework for Adaptive Collaborative Perception in IIoT Robotic Swarms
- Execution Flexibility in Automated Planning: A Comparative Evaluation of Deordering and Reordering Strategies
- SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning
- ORDER: Task-Conditioned Routing for Retrieval-Augmented Generation
- FlexEE: Self-Speculative and KV-Compatible Early Exiting for Offloading-Aware LLM Inference
- Scaling-Score Conformal Prediction for Multi-Target Regression
- Sample-Conditioned Representation Selection for Audio Few-Shot Learning
- ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds
- FirmCORe: A Benchmark for Structured Reasoning about Inter-Firm Collaboration Opportunities
- Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs
- Semi-Supervised Learning-Based Genetic Biomarkers Dataset for Multiple-Stage Hepatocellular Carcinoma Prediction
- ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training
- EchoPath: Execution-Level Replayable Memory for GUI Agents
- Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling
- From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts
- Extracting ontology-compliant knowledge from scientific text describing irradiated materials using large language models
- Talking Head Synthesis with Facial Landmark Guidance via 3D Gaussian Splatting
- Transformer-Based Token Fusion and Dynamic Graph Planning for Audio-Visual Navigation
- World Model Science: Self-Organized Criticality, Weak Chaos, and Metastable Belief Dynamics in Long-Horizon LLM Agents
- LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
- Verifiable Social Reasoning for LLM Assistants
- ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
- Sparse MLLM Anchors, Dense Adaptation: Breaking the Self-Referential Loop in Wild Test-Time Adaptation
- ThinkFlow: Self-Evolving Probabilistic Latent Memory for Lifelong Conversational Agents
- Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition
- AntennaFlow: A Generative Flow Model for Offset Correction in Phaseless Antenna Testing
- Bridging Learned Visual Perception and Symbolic Belief-Space Planning
- AI for Games in the Foundation Model Era
- A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance
- From Manual Construction to AI-Driven Scenario Emergence: Rethinking Catastrophe Risk Modeling
- Toward Governance-Aware Autonomous GIS: A Narrative Review of Ethical and Privacy Risks in LLM-Enabled GeoAI
- JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management
- Never Stop Thinking: Continuous-Time Language Agents
- FlashVector: Agent for Hierarchical Model Serving Stack Optimization
- Intrinsic Motivation in Reinforcement Learning: A Research Agenda for Adaptive Self-Organisation
- Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?
- Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models
- Coverage-Aware Virtual IMU Augmentation for Low-Resource Human Activity Recognition
- little m: An AI Agent for Industrial Process Optimization
- GPEvac: GNN-Based PPO for Adaptive Evacuation Routing During Shooting Events
- Neuro-Symbolic Hierarchical Intention Anticipation in Human Behavior
- Skill-based Agentic Evaluation for Real-time Data Science Tasks
- BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
- Closing the Loop: Branch-and-Bound for Scalable Verification of Nonlinear Neural Feedback Systems
- End-to-End Latency-Minimizing and Load-Balanced Request Scheduling for Edge LLM Inference in Agentic AI Services
- MOCC-R1: Reinforcing Reasoning-Response Consistency for Multimodal Counselor Response Generation
- Symbolic Separation: Grounding Deep Agents in Knowledge Graphs for Trustworthy Operational Data Analytics
- Interactive Memory Learning for Long-Term Conversations
Comments
Please log in to post a comment.