Researchers have proposed various methods to improve the performance of large language models (LLMs) in time series forecasting, including the use of retrieval-augmented generation (RAG) and the development of novel architectures such as TS-RAG. These methods have shown promising results in improving the accuracy of LLMs in time series forecasting tasks. However, the application of RAG in time series forecasting remains limited, and most time series models are constrained by limited training data, smaller parameter scales, and a lack of extensive generative capabilities. To address these challenges, a novel approach, TS-RAG, has been proposed, which leverages RAG to enhance forecasting performance. The framework introduces specially designed reference tokens to effectively fuse information from the input sequence with that from retrieved similar sequences, enabling a more robust capture of complex temporal dynamics. Experimental results demonstrate that TS-RAG achieves consistent state-of-the-art performance across several real-world forecasting benchmarks.
A hybrid machine learning framework has been developed for herd-level cattle growth pattern and weight gain forecasting in grazing-based production systems. The framework integrates weekly live weight observations, demographic variables, and lagged environmental predictors into structured forecasting datasets. Herd-level forecasting trajectories were generated through temporal aggregation of animal-level predictions. Four hybrid architecture families were evaluated, including residual, stacked, cascade, and ensemble-assisted frameworks. ARIMA, LSTM, and GRU models were used as comparative baselines. Independent testing demonstrated strong predictive agreement across multiple forecasting horizons. The cascade GB to RF to NN architecture achieved the best performance, with a test R^2 of 0.889, RMSE of 21.319 kg, and MAE of 15.462 kg. Hybrid architectures maintained greater robustness than recurrent sequential models under sparse observation conditions.
A controlled multimodal benchmark, C-SUITEBENCH, has been introduced to evaluate the decision-making ability of large language models in executive business decisions. The benchmark includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. The results show that multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, a multimodal integration paradox was uncovered, where adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves.
A benchmark revision approach has been proposed to improve the realism of synthetic clinical benchmarks under utility constraints. The approach is formulated as utility-constrained realism improvement, where dataset changes should increase realism while staying above an operational utility floor. The results show that two deterministic revisions improve the realism of the benchmark while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating.
A hybrid machine learning framework has been developed for herd-level cattle growth pattern and weight gain forecasting in grazing-based production systems. The framework integrates weekly live weight observations, demographic variables, and lagged environmental predictors into structured forecasting datasets. Herd-level forecasting trajectories were generated through temporal aggregation of animal-level predictions. Four hybrid architecture families were evaluated, including residual, stacked, cascade, and ensemble-assisted frameworks. ARIMA, LSTM, and GRU models were used as comparative baselines. Independent testing demonstrated strong predictive agreement across multiple forecasting horizons.
A novel approach, TS-RAG, has been proposed to leverage retrieval-augmented generation (RAG) to enhance forecasting performance in time series forecasting tasks. The framework introduces specially designed reference tokens to effectively fuse information from the input sequence with that from retrieved similar sequences, enabling a more robust capture of complex temporal dynamics. Experimental results demonstrate that TS-RAG achieves consistent state-of-the-art performance across several real-world forecasting benchmarks.
A hybrid machine learning framework has been developed for herd-level cattle growth pattern and weight gain forecasting in grazing-based production systems. The framework integrates weekly live weight observations, demographic variables, and lagged environmental predictors into structured forecasting datasets. Herd-level forecasting trajectories were generated through temporal aggregation of animal-level predictions. Four hybrid architecture families were evaluated, including residual, stacked, cascade, and ensemble-assisted frameworks. ARIMA, LSTM, and GRU models were used as comparative baselines.
A controlled multimodal benchmark, C-SUITEBENCH, has been introduced to evaluate the decision-making ability of large language models in executive business decisions. The benchmark includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. The results show that multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification.
A benchmark revision approach has been proposed to improve the realism of synthetic clinical benchmarks under utility constraints. The approach is formulated as utility-constrained realism improvement, where dataset changes should increase realism while staying above an operational utility floor. The results show that two deterministic revisions improve the realism of the benchmark while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating.
A novel approach, TS-RAG, has been proposed to leverage retrieval-augmented generation (RAG) to enhance forecasting performance in time series forecasting tasks. The framework introduces specially designed reference tokens to effectively fuse information from the input sequence with that from retrieved similar sequences, enabling a more robust capture of complex temporal dynamics. Experimental results demonstrate that TS-RAG achieves consistent state-of-the-art performance across several real-world forecasting benchmarks.
Key Takeaways
- Researchers have proposed various methods to improve the performance of large language models (LLMs) in time series forecasting, including the use of retrieval-augmented generation (RAG) and the development of novel architectures such as TS-RAG.
- A hybrid machine learning framework has been developed for herd-level cattle growth pattern and weight gain forecasting in grazing-based production systems.
- A novel approach, TS-RAG, has been proposed to leverage retrieval-augmented generation (RAG) to enhance forecasting performance in time series forecasting tasks.
- A controlled multimodal benchmark, C-SUITEBENCH, has been introduced to evaluate the decision-making ability of large language models in executive business decisions.
- A benchmark revision approach has been proposed to improve the realism of synthetic clinical benchmarks under utility constraints.
- A hybrid machine learning framework has been developed for herd-level cattle growth pattern and weight gain forecasting in grazing-based production systems.
- A novel approach, TS-RAG, has been proposed to leverage retrieval-augmented generation (RAG) to enhance forecasting performance in time series forecasting tasks.
- A controlled multimodal benchmark, C-SUITEBENCH, has been introduced to evaluate the decision-making ability of large language models in executive business decisions.
- A benchmark revision approach has been proposed to improve the realism of synthetic clinical benchmarks under utility constraints.
- A novel approach, TS-RAG, has been proposed to leverage retrieval-augmented generation (RAG) to enhance forecasting performance in time series forecasting tasks.
Sources
- TS-RAG: Retrieval Augmented Generation for Time Series Forecasting
- Hybrid Machine Learning Framework for Herd-Level Cattle Growth Pattern and Weight Gain Forecasting in Grazing-Based Production Systems
- Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
- Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
- HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
- Abstract Event Causal Rules: Induction and Application
- SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
- ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution
- OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality
- Posture and Sustainment Optimization Under Adversarial Uncertainty
- Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
- DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data
- Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services
- SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse
- From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction
- Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
- ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
- Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows
- StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
- EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
- Comparative Approaches to Agent Retrieval over Large Skill Libraries
- VLMs for Videogame Data Annotation
- MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration
- Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI
- RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
- DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
- Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks
- When Agentic AI Meets Integrated Sensing and Communication
- Cautious Context Steering for Language Model Personalization
- The Ignition Index: Measuring Global Workspace Dynamics in Language Models
- When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
- Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index
- Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing over Prerequisite DAGs
- Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation
- When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment
- Subliminal Learning is Non-Semantic Distillation
- SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation
- Shaping Human-AI Interactions to Provide Improvement Pathways and Balance Competing Objectives
- OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents
- AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
- Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment
- Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process
- Training a Conditioned Video Game Agent on a VLM Annotated Dataset
- TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
- Challenges in Evaluating Explanation Methods for Static and Evolving Data
- Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
- The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
- Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
- Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
- QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction
- PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
- The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
- Runtime Observability for Heterogeneous Attention Memory
- AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
- ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation
- When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
- EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
- Unified Agent: Managing Interactions across Devices
- BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming
- Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks
- Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging
- Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model
- Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
- When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
- Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis
- From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
- HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
- Contextual Information Policy Optimization for Search Agents
- Mind the Gaps: Mixture-of-Minds for Human Simulation
- From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems
- WorldClaw: Agentic 3D Open-World Generation at Scale
- CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
- iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
- FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
- Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts
- ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment
- Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference
- GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
- CourseGraph: Finding overlaps and differences in Computer Science courses across universities
- GSBF: Gaussian Splatting for Environment-Aware Beamforming
- Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding
- SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
- Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability
- Recursive Synthesis for Long-Horizon Terminal Tasks
- SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications
- LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
- Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning
- Coherence-Oriented Dream Scene Visualisation
- TriQua: Reconciling Granularity and Context in Factuality Evaluation
- Project2Task: Graph-Guided Project-Level Planning for Autonomous Research
- Small Foundation Models of Human Cognition and Behaviour
- PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
- Otter: A Time-Aware, History-Conditioned Human Chess AI
- Measuring and Detecting Harmful AI Sycophancy
- Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Plan Coordination
- C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models
- Counterfactual Analysis via Large Language Models
- CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction
- DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
- A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition
- Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
- Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying
Comments
Please log in to post a comment.