Harness Compilation Boosts VLM Scores by 9.9-23.9 Points, LLoCoT Reduces Latency by ~36x

Recent advances in LLM efficiency and reliability include Harness Compilation, which improves VLM scores by 9.9-23.9 points, and LLoCoT, a looped latent-reasoning framework reducing latency by ~36x. Plan-and-Patch demonstrates diffusion models outperform autoregressive ones in plan repair (53.7% vs 27.0%) while cutting latency by 39-46%. RaReCache enables cross-model KV cache reuse, retaining 95-99% accuracy across large parameter gaps, and Memento 3 allows frozen agents to learn world models, clearing all 25 ARC-AGI-3 games with 100% RHAE. OnTrack detects failures in ~1ms to save 18% compute, and EvoAlloc reduces evaluation costs by 59-82% via self-evolving resource allocation.

Agent safety and reliability remain critical, with typed decision models showing unreliable accuracy (36%-72%) against defenses like prompt injection. Legal reasoning models exhibit disconnects between citations and verdict dependence, while HarmBench fails to measure single 'harmful refusal' attributes due to saturation. Safety studies reveal 56.92% of trajectories contain unfulfilled obligations, prompting ObligationGuard with 56.52% recall. Ecological safety theory suggests misaligned populations face critical takeoff thresholds, and BrickBench finds agents satisfy physical constraints but fall short of human design quality in LEGO tasks.

Methodological innovations address calibration and reasoning gaps. Bayesian neural networks achieve efficient inference by dynamically allocating Monte Carlo samples, reducing latency while maintaining error guarantees. TRACE traces emotions in social scenes to reveal model gaps, and ReCast improves failure attribution by learning step representations for best Hit@1 performance. Universal Textual Teaching raises student accuracy on math and code tasks via parameter-free textual primers, and GeoReform boosts geometry reasoning accuracy from 42% to 56% by treating formalization as an optimizable policy.

Domain-specific applications show significant progress. Synthetic vascular audits predict equivalence within 4.0e-15, though noisy anchors constrain calibration. OA-MAP achieves 0.80 AUROC for knee osteoarthritis structural progression using multimodal data. AtomWorld-Mirror accelerates materials simulation 1000-10000x via macro-step inference, and BEVPIPE enables portable BEV perception deployment with 19.5x speedup. TokenBank uses structured contracts to manage AI service costs, reducing mean expenditure by $304.88, while Agent-controlled forgetting cuts token usage by 50% and costs by ~67% in tool-use scenarios.

Key Takeaways

  • Harness Compilation improves VLM scores by 9.9-23.9 points.
  • LLoCoT reduces latency by ~36x via looped latent reasoning.
  • Diffusion models outperform autoregressive ones in plan repair (53.7% vs 27.0%).
  • RaReCache retains 95-99% accuracy across large parameter gaps.
  • Memento 3 clears all 25 ARC-AGI-3 games with 100% RHAE.
  • Typed decision models show unreliable accuracy (36%-72%) against defenses.
  • HarmBench fails to measure single 'harmful refusal' attributes due to saturation.
  • GeoReform boosts geometry reasoning accuracy from 42% to 56%.
  • AtomWorld-Mirror accelerates materials simulation 1000-10000x.
  • Agent-controlled forgetting cuts token usage by 50% and costs by ~67%.

Sources

NOTE:

This news brief was generated using AI technology (including, but not limited to, Google Gemini API, Llama, Grok, and Mistral) from aggregated news articles, with minimal to no human editing/review. It is provided for informational purposes only and may contain inaccuracies or biases. This is not financial, investment, or professional advice. If you have any questions or concerns, please verify all information with the linked original articles in the Sources section below.

ai-research machine-learning arxiv research-paper harness-compilation llcot plan-and-patch rarecache memento-3 agent-controlled-forgetting

Comments

Loading...