Research Brief
Key Takeaways
- • Key findings from research papers
Sources
- Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System
- Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps
- PLeDO: Pain Level Detection for Osteoarthritis from EMR Data
- HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation
- RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning
- Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State
- Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration
- The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale
- Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation
- CoupVisor: Strategy Optimization by Round and Challenge Decision Support
- RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration
- Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning
- Breaking and Defending LLM-Powered Social Media Bot Detection Systems
- Augmenting Text to Increase Translation Difficulty
- UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
- Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics
- Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
- MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment
- ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction
- When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification
- Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling
- FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
- Competing at Every Price Point with Agentic Evolution over a Menu of LLMs
- Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior
- Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication
- AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems
- DriveCache: Action-Aware Caching for Driving World Model Inference
- AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
- Process-Constituted Intelligence: A Shared Criterion for Humans and Machines
- ParaTempo: Efficient Parallel Reasoning via Temporal Confidence
- A Policy Algebra for Trust-Preserving Agentic AI Execution
- Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics
- The Value of a Prompt: An LLM-Relative Kolmogorov-Complexity Approach
- Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation
- JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills
- CUBICS: Situation-aware performance estimation for safety-relevant ML components
- CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction
- Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
- Probabilistic Circuits as Reasoning Machines in Artificial Intelligence (Part I)
- FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
- Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement
- A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory
- Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
- PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning
- Quipu: A Governed Bitemporal Knowledge Graph Store
- When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding
- LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing
- What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
- Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
- Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152
- What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics
- Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
- GRIP: Grounded Reasoning via Information-Restricted Premises
- Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents
- Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling
- HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents
- Drive, Pack, Fly: The Travelling Thief Problem with Drone
- Reasoning-supported Robustness Validation of Automotive E/E Components
- Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity
- Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance
- Solvable Sokoban Without a Solver via Diffusion
- Navigation-Informed Embeddings: Dense-Retriever Adaptation from Agent Search Traces
- KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving
- Intent-Driven Situation Tracking for User-Centric Multi-Turn Agents
- Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign
- Adaptive Mixing of Policies from Searching and Policies from Learning
- From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM
- From Contexts to Values: Context-Dependent Defeat in Abstract Argumentation
- Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation
- A Network-driven Framework for Public Event Forecasting via Dynamic Interaction Network Evolution
- Mental Model Management: An Operator-Based Framework for LLM Memory
- Incoherent by Design? On the Moral Self-Consistency of LLMs
- The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
- Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis
- ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits
- Validation-Frontier Representation Selection under Constrained Observation
- Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning
- Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts
- Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World
- Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy
- GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG
- Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid
- Personalized Auto-Research: Towards a True AI Co-Scientist
- JarvisBench: Always-on Intelligence Between Humans and Agents
- MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
- Synchronized Logit Steering: Real-world Steganography
- Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
- When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
- Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry
- SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization
- Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study
- Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration
- Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate
- DeepInsight II: One Trace from Benchmark to Robot
- Bounded Agents: Delegation Security for Multi-Agent AI Systems
- Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs
- A survey of AI-generated voices and their detection
- Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL
- Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
- Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations
- Learning Agent Execution for KV-Cache Management in Agentic Serving
- Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
- FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment
- Position: AI Lock-In Is in Progress, and We Must Be Prepared
- Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review
- When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL
- Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws
- From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change
- Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture
- OGX: An Open-Source, Vendor-Neutral Generative AI Application Server
- A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
- The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines
- An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case
- When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning
- Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
- Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring
- Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
- Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
- A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
- When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
- Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma in Quantum Parallel Repetition
- Beyond Correctness: Toward Automated Novelty Verification with Lean 4
- Advanced modelling and data analytics in aviation
- Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
- Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous
- CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
- From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
- Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
- What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
- Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
- When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
- Small Models Scout Bottleneck Order for Large-Model Data Control
- LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks
- Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
- Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
- S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
- Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5
- T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework
- TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning
- Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research
- LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning
- Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems
- Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
- LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents
- Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark
- Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads
- Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents
- ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models
- LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
- SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion
- ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning
- Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark
- $D^{2}R^{2}$: Discrete Diffusion with Regulation Reinforcement for Single-Cell Perturbation Prediction
- VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
- Decentralized Federated Learning for Heterogeneous Multi-Task Semantic Communication
- Physics-informed VAE-EVT for Tail Aware Radio Map Prediction
- MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning
- Physiological World Models for Human State Transitions
- Understanding Cognition-Induced Risks in Agentic AI Systems
- UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms
- A concentration result for multilayer feedforward neural networks
- Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
- FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA
- Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs
- Large Language Model Assisted Operational Monitoring for Battery Energy Storage System Integrated Power Distribution Networks
- Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish
- TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions
- Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization
- Does the Proof Prove It That Way? Faithful Formalization of Elements Proofs
- EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints
- Dynamic Multi-Byte Prediction With Hierarchical Language Models
- Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
- ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search
- When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction
- Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback
- Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds
- TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation
- Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
- THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts
- A Responsible Artificial Intelligence Framework for Groundwater Modeling
- Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict
- Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs
- Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
- RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder
- Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
- Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
- Large Language Models and their Awareness of Mechanics and Spatial Geometry
- Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP
- Position: Medical AI Neglects Real Treatment Outcomes
- BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics
- Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain
- TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
- OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks
- SCOPE: Score-Isolated Agentic Optimization for Video World Models
- Demand-Driven Vertiport Siting and Discrete-Event Fleet Simulation for On-Demand Urban Air Mobility Network Design
- The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
- Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation Shifts
- VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
Comments
Please log in to post a comment.