Skip to content
Monday, July 20, 2026
Research

Research

Research coverage with in-house analysis.

Research

CoCoDA Lets Small Models Grow a Living Tool Library

Researchers propose CoCoDA, a framework that co-evolves a language model planner and its tool library via a compositional code DAG. Typed retrieval prunes candidates symbolically, keeping context costs fixed as the library grows.

Source: Read full story at arXiv.org

In-house analysis

Capability without context discipline is a budget problem. The real test is whether sublinear retrieval survives messy, open-domain tool libraries beyond controlled benchmarks.

Research

Researchers Propose Game Theory Fix for Sycophantic AI

Researchers formalize AI sycophancy as a Crawford-Sobel cheap talk game, where chatbots reinforce both truth-seekers and validation-seekers identically. A proposed Epistemic Mediator intervention achieved a 48x differential in belief-spiral rates in simulation.

Source: Read full story at arXiv.org

In-house analysis

The tension here is satisfaction versus accuracy. Operators optimized for engagement metrics will resist friction by design, which is precisely where the spiral starts.

Research

LLMs Shift Political Stance When Users Ask Them To

Researchers tested 200 politically oriented questions across economic and personal freedom axes. User prompts reliably shifted responses in larger, newer models; system prompts largely failed to do so.

Source: Read full story at arXiv.org

In-house analysis

Capability is not neutrality. The models most trusted for sensitive work are also the most politically steerable, which is the exact contrast buyers and regulators need to price into deployment decisions.

Research

SkillLens Cuts LLM Agent Costs With Hierarchical Skill Reuse

Researchers propose SkillLens, a four-layer skill graph that lets LLM agents reuse only the relevant parts of past experience. On ALFWorld, success rate rose from 45.00% to 51.31%, with a 6.31 percentage-point accuracy gain on bug localization.

Source: Read full story at arXiv.org

In-house analysis

Capability is cheap. Selective reuse is the cost control. Operators who route agent memory this precisely pay for adaptation, not repetition.

Research

Grid Overlays Beat Semantic Prompts for Chart Data Extraction

Researchers found that overlaying a coordinate grid on chart images reduced extraction error significantly, from 25.5% to 19.5% SMAPE. Chain-of-Thought and metadata-first semantic methods produced no statistically significant improvement.

Source: Read full story at arXiv.org

In-house analysis

Semantic guidance is expensive to maintain. A coordinate grid costs nothing. For operators building literature-analysis pipelines, input preprocessing beats prompt engineering here.

Research

SFT vs RL Is the Wrong Question in Post-Training

Researchers argue post-training debates misplace the key distinction. What matters is whether training expands a model's reachable behaviors, not whether the method is SFT or RL.

Source: Read full story at arXiv.org

In-house analysis

Elicitation versus creation is the contrast that now governs post-training investment decisions. Labs that conflate the two will misread what their fine-tuning actually buys.

Research

Standard Text Embeddings Fail to Capture Human Preferences

Researchers show standard text embeddings measure semantic similarity, not preferential agreement, making them unreliable for collective decision-making. Synthetic training data that breaks the semantic-preference correlation improved preference prediction across 11 deliberation datasets.

Source: Read full story at arXiv.org

In-house analysis

The gap between semantic similarity and preferential agreement is the gap between what a model reads and what a person means. Operators building on standard embeddings are buying the wrong plumbing.

Research

Attention Maps Are Useless Predictors of VLM Correctness

Researchers tested three open-weight VLMs and found attention structure predicts correctness at near zero (R=0.001). Hidden-state probes and self-consistency at K=10 are far stronger reliability signals.

Source: Read full story at arXiv.org

In-house analysis

Builders who filter outputs by attention confidence are watching the wrong signal. Hidden-state probes work, but architectural fragility, concentrated versus distributed, determines how much that monitor can be trusted.

Research

One Prompt Sentence Makes AI Models Significantly More Creative

Researchers found that adding a single sentence to an AI prompt measurably increases model creativity. The finding requires no new model, no fine-tuning, only a change to the input.

Source: Read full story at VentureBeat

In-house analysis

Capability is not the constraint here. Prompt craft is. The operator who ships this into production tomorrow beats the lab that trains for six more months.

Research

AI Companion Market Hits $381B With Documented Child Harms

The AI companion market is projected to reach $381 billion by 2032. Meanwhile, lawsuits against CharacterAI and internal Meta documents reveal serious harm risks, especially for minors.

Source: Read full story at Tech Policy Press

In-house analysis

The market scales faster than the safeguards. The signal to watch is not capability growth but who absorbs liability when a $381 billion industry builds on documented, categorized child harms.