Recent research advances AI efficiency, safety, and reliability across diverse domains. Jev-style models exhibit 'option-label bias' while PACE reduces token usage by 17-25%; GraphMemory and MRVQ cut token and memory consumption by 81-85% and 17.8-22.0x respectively. TradeGrad optimizes trading strategies, JOVE reduces cost/latency by 3.17x, and TACD achieves 7.7-11.9x speedups. In healthcare, HEAR lowers heart rate error to 1.6 BPM, FD-SCoPE answers clinician questions with 90.7-97.9% accuracy, and Breast cancer AI reaches 0.79 AUROC. THPL boosts aquaculture decision accuracy from 33% to 97%, and SoftGene enhances gene annotation.
Safety auditing via Positive-Unlabeled learning outperforms baselines by 5-17 points, and HACKTRACE detects reward hacking with 0.997 AUC. LUMOS reveals a retrieval gap where 84% of internal encoding fails to translate to 54% behavioral expression. Law&Order achieves 100% tax law autoformalization accuracy, yet LLM judges fail to match worker acceptance rates (3.0%-97.9% vs 61.1%). Jailbreak Benchmark reports unsafe rates rising to 18.65%, while PlurPO cuts endorsement of harmful actions by 89%. Algorithmic extraterritoriality describes AI-driven financial surveillance, and UK AI risk disclosures rose to 41.2% but substantive ones remain rare at 4.3%.
Efficient reasoning reduces CoT faithfulness but preserves monitorability, whereas Loop models may reduce CoT monitorability under stress. Dynamic routers underperform due to 'difficulty blindness', and Near-Zero Monitor Readout does not confirm behavioral control. Ego2World improves action validity by 4.15 points, and VIGOR achieves 43.6% gain on Robosuite. Dual-Stream OSSE-LSTM achieves 96.36-96.72% accuracy, and MetaRubric improves medical QA accuracy by 20.40 points. FlashSinkhorn 2 solves discrete EOT for 1.34×10⁸ particles in under 2.5 hours, and CALM improves text-to-image safety. GHOST hazards occur in 11.5% of GPT-5.5 interactions, and CreateScore reduces token use by 65.2% though it misses 8B model errors.
Key Takeaways
- PACE reduces token usage by 17-25%.
- GraphMemory uses 81-85% fewer tokens.
- MRVQ uses 17.8-22.0x less memory.
- JOVE reduces cost/latency by 3.17x.
- HEAR reduces HR error to 1.6 BPM.
- FD-SCoPE answers clinician questions with 90.7-97.9% accuracy.
- Law&Order achieves 100% tax law autoformalization accuracy.
- PlurPO cuts endorsement of harmful actions by 89%.
- FlashSinkhorn 2 solves discrete EOT for 1.34×10⁸ particles in <2.5 hours.
- GHOST hazards occur in 11.5% of GPT-5.5 interactions.
Sources
- Labels Override Definitions in Jev-Style Typed Decision Models
- Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
- When to Compile a Computer-Use Agent? Measuring Payback and Making Compilation Decisions for Token Efficiency
- Positive-Unlabeled Learning for Agent Safety False Alarm Auditing
- HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence
- Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM
- LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs
- Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability
- THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS
- Decoupling Memory from Context: Structured Memory for Token-Efficient Test-Time Continual Learning
- Trading Strategy Optimization via Textual Gradient
- Learning to Assess Heartbeat Observability for mmWave Heart-Rate Sensing
- NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents
- Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System
- Transcriptome-informed multi-modal AI for predicting neoadjuvant therapy response from breast cancer biopsies
- MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search
- JOVE: Joint Execution and Verification for Resource-Aware LLM Task Graphs
- EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation
- hacktrace: behavior-supervised detection of reward hacking during code generation
- Becoming Suspicious Across Borders: Algorithmic Extraterritoriality and AI-Driven Financial Surveillance
- When Numbers Start Talking: Numerical Signalling and Strategic Behaviour Among LLMs
- SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation
- "I just assumed that it would translate": examining MT risk awareness among healthcare staff with abbreviations as a use case
- Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
- The AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?
- PAPER2LLM++: Continual Self-Evolution of LLMs from Research Papers
- Law And Order: Tax Law Autoformalization
- Dynamic LLM Routers are Often Misguided
- On the Chain-of-Thought Monitorability of Looped Language Models
- Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
- Lost in the Request: How Communication Variation Disrupts Retrieval and Action in Email Agents
- Open-Endedness Bench: Measuring Epistemic Process from Agent Records
- Answering clinicians' questions over trial evidence tables with verifiable, feedback-driven language models
- FlashSinkhorn 2: Block-Sparse Entropic Optimal Transport
- KV$^2$: A Self-Refining KV Cache
- Interpreting at Write Time: A Policy Ablation for Multi-Goal Agent Memory
- iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
- Learning a Fact Is Not Learning How to Retrieve It
- Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
- How to Have a Sensitive Debate: An Instance-Optimal Protocol for AI Debate
- Bounded Reachability & Jailbreak Detection via Contraction-Constrained State Space Models
- PsyEvo: A Personalized Counseling Agent That Self-Evolves at Test Time
- DyadMem: A Long-Term Memory Benchmark of How Agents Work with Users
- RIFAR: Reliability and Forgetting-Aware Replay for Continual Robot Learning
- ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models
- CVE2AP: Automated Generation of PDDL-Encoded Attack Paths via Large Language Models
- Verifiable, Articulable, and Tacit Components of Preference
- Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations
- RASPER: Reward-Aligned Summarization of Clinical Notes for EHR Outcome Prediction
- Relevant Evidence Decoding for Audio-Visual Hallucination Mitigation
- TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control
- Harness-Aware Distillation for Small Language Model Agents
- DNAlign: Dynamic Null-Space Safe Alignment for LLMs
- AMBER: Multi-View Adaptive Budget Allocation for Listwise Vision-Language Reranking
- FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training
- MLCommons Jailbreak Benchmark v1.0
- Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
- MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning
- Label-Efficient Time Series Classification at Scale: A Dual-Stream OSSE-LSTM with Counterfactual Attribution
- Learning to Revise Reasoning with Segment-wise On-Policy Distillation
- Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves
- DataWeave: Deploying Human-LLM Analytics for Exploratory Structured Data Analysis
- Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation
- MintFlow: Minimal Trajectory Intervention for Constrained Flow Matching
- Knowledge or Calculator? Decomposing the Skill Premium in Verifiable Financial Agent Workflows
- Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
- From Benchmarks to Production: A Text-to-SQL System for Complex Financial Data
- Reasoning Models Are Accurate but Unsound on Identification
- Keeping JEPA World Models Plannable When Little of the Frame Moves
- Predictor-Guided Latent Space Codon Optimization for Maximizing Protein Expression
- Peer Influence across Heterogeneous AI Models
- MOF-VERIFY: A Failure-Aware Agentic Harness for MOF Hypothesis Verification
- Reliable Self-Evolution with Imperfect Proxy Rewards
- Reasoning with Evidence, Not Merely Rationales: Verifiable Preference Proofs for LLM-Based Recommendation
- Continual Graph Memory for Mathematical Research Agents
- CuBEs: Culturally-Situated Behavioral Evaluations and the Limitations of Culture-Blind LLM Judges
- VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
- Time Series Forecasting Benchmarks Need Scenario-Grounded Stress Testing
- TasteBench: Multimodal Benchmark for Sensory Prediction, from Molecules to Sustainable Foods
- Mitigating Social Sycophancy via Pluralistic Preference Optimization
- Reinforcement Learning Techniques for the Optimization of Target Polarization in Nuclear Physics Scattering Experiments
- When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge
- ROUTEAUDIT: Interaction-Aware Identification for Budgeted Multi-Verifier Routing
- A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control
- Jumping the Line: Exploiting Length Predictions in LLM Scheduling
- Benchmarking Candidate Coverage in Typed Decision Models
- MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations
- Tropical Reinforcement Learning
- Toward SLM-based agentic task-tool intent matching
- VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning
- BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration
- On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor
- World Action Modeling with Progressive Visual Planning
- Modeling Shared and Individual Structure for Cross-Subject Continuous Affect Regression from EEG-fNIRS
- CreateScore: Domain-Theory-Informed Bayesian Routing for LLM-Based CV Screening
- PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation
- Safeguarding Mutual Correction in Source-Free Domain Adaptation via Cut Statistics
- Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
- Depth as Time in One-Step Generative Models
- Low-Cost Video--Time Priors as a Strong Baseline for EEG--fNIRS Emotion Regression on Familiar Videos
- HazardWeaver: Scientific Route Selection for Hazard Analysis Agents
- HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
- Multilingual GSM-Symbolic: What determines capability transfer across languages?
- Geometry Meets Physics: Data-Efficient Pre-Training for Unstructured Neural PDE Solvers
- Preserving Mathematical Reasoning in Compressed Diffusion Language Models via Trajectory-Aware Low-Rank Approximation
- Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS
- Multi-Task Evolution for Zero-Shot Cross-Problem Generalization using LLMs
- Lightweight, Rubric-Guided Trajectory Evaluation for Production AI Agents
- Optimal Planning in a Dynamic World
- A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns
- Coherence-Driven Belief Formation and Population Dynamics of Contagion in LLM Agents
- Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen LLM Hear Beyond the Transcript
- Designing the Future of User Feedback for Generative AI
- How To Train Your World Model: Fine-tuning vs RAG for LM-based World Modeling
- Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
- Hypothesis-guided discovery of cognitive algorithms via program refinement
- HXAI: Hierarchical Privacy-Preserving Explainable AI in Distributed Energy Systems
- What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute
- Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation
- Traversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion
- DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
- A Multi Method Importance and Performance Efficiency Analysis of Topological Metrics for Natural Visibility Graph Based Cyber Attack Detection
- World Editing: Intervening on Executable Worlds at Increasing Depth
- Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
Comments
Please log in to post a comment.