Skip to content
Sunday, July 19, 2026

The essential AI briefing, no fluff

Research||arXiv.org|Paper

CoCoDA Lets Small Models Grow a Living Tool Library

Research illustration

Published on arXiv, CoCoDA addresses a scaling failure in tool-augmented language models. As tool libraries grow, retrieval costs balloon and prompt budgets break. The framework structures tools as nodes in a directed acyclic graph, with typed signatures, behavioral specs, and worked examples attached to each node.

Continue reading

Researchers propose CoCoDA, a framework that co-evolves a language model planner and its tool library via a compositional code DAG. Typed retrieval prunes candidates symbolically, keeping context costs fixed as the library grows.

In-house analysis

Capability without context discipline is a budget problem. The real test is whether sublinear retrieval survives messy, open-domain tool libraries beyond controlled benchmarks.

Read our brief

Latest

See more
Research||arXiv.org|Paper

Grid Overlays Beat Semantic Prompts for Chart Data Extraction

Research illustration

A paper published on arXiv tested two competing strategies for getting multimodal LLMs to extract data from scientific charts: high-level semantic prompting versus low-level spatial priming. Spatial priming won. Overlaying a coordinate grid on the chart image before analysis reduced SMAPE from 25.5% to 19.5%, a statistically significant result (p < 0.05).

Continue reading

Researchers found that overlaying a coordinate grid on chart images reduced extraction error significantly, from 25.5% to 19.5% SMAPE. Chain-of-Thought and metadata-first semantic methods produced no statistically significant improvement.

In-house analysis

Semantic guidance is expensive to maintain. A coordinate grid costs nothing. For operators building literature-analysis pipelines, input preprocessing beats prompt engineering here.

Read our brief

Models

See more
Models||Anthropic

Anthropic Details Fable 5 Cyber Safeguards and a Jailbreak Severity Framework

Anthropic Details Fable 5 Cyber Safeguards and a Jailbreak Severity Framework

Anthropic has published a detailed account of the cyber safeguards shipping with Fable 5, as the model returns to global availability. The post describes safety classifiers tuned with a larger margin than previous models, sorting cybersecurity requests into four categories: prohibited use, high-risk dual use, low-risk dual use, and benign activity. The stated aim is to allow defensive security work while blocking tasks that enable attacks.

Continue reading

Anthropic published the technical detail behind Fable 5's cybersecurity guardrails: safety classifiers that sort requests into prohibited, high-risk dual use, low-risk dual use, and benign categories, plus a draft Cyber Jailbreak Severity scale and a HackerOne program for responsible disclosure.

In-house analysis

The framework is the story, not the classifiers. Whoever writes the shared severity language for AI jailbreaks sets the terms regulators later adopt, and Anthropic just published the first draft.

Read our brief

Policy

See more
Policy||GOUV

France Calls for Urgent AI Regulation on Three Pillars

Policy illustration

France's Ministry for Europe and Foreign Affairs is calling for binding AI regulation, framing the absence of oversight as an active risk. The ministry cites three founding concerns: investment and innovation, digital sovereignty, and ethical and human considerations. The European Commission's February 2020 White Paper on Artificial Intelligence is named as the policy anchor.

Continue reading

France's foreign ministry is calling for urgent AI regulation built on three principles: investment and innovation, digital sovereignty, and ethical considerations. The push references the European Commission's February 2020 White Paper on Artificial Intelligence.

In-house analysis

The contrast is clear: private labs own the capability, but regulators will own the accountability rules. Who pays for compliance is the next question.

Read our brief

Research

See more
Research||arXiv.org|Paper

LLMs Shift Political Stance When Users Ask Them To

Research illustration

A new arXiv paper introduces the concept of political plasticity: how readily a model shifts its ideological stance based on supplied context. The researchers tested 200 questions across economic and personal freedom axes, drawing on a framework from Lester (1996). User prompts with few-shot examples produced significant ideological shifts in larger, newer frontier models. System prompts were largely ineffective.

Continue reading

Researchers tested 200 politically oriented questions across economic and personal freedom axes. User prompts reliably shifted responses in larger, newer models; system prompts largely failed to do so.

In-house analysis

Capability is not neutrality. The models most trusted for sensitive work are also the most politically steerable, which is the exact contrast buyers and regulators need to price into deployment decisions.

Read our brief

Industry

See more
Industry||CRN

SAP Buys Dremio and Prior Labs for Agentic AI Push

SAP Buys Dremio and Prior Labs for Agentic AI Push

SAP is acquiring Dremio and Prior Labs, pairing the deals with a $1.1 billion investment commitment to scale Prior Labs globally. Dremio brings Apache Iceberg-native lakehouse architecture. Prior Labs brings Tabular Foundation Model research. Together they slot into SAP Business Data Cloud.

Continue reading

SAP is acquiring data lakehouse provider Dremio and AI model specialist Prior Labs. A $1.1 billion investment commitment will scale Prior Labs globally to advance Tabular Foundation Models.

In-house analysis

SAP is buying the plumbing so agents have clean water to run on. The operator question: does owning the data layer convert to owned AI outcomes, or just higher maintenance costs?

Read our brief

AI Use Cases

See more

German researchers are using AI to process data from more than 1,000 camera traps, sound recorders, and climate sensors across national parks. The system identifies animal species, tracks populations, and analyzes forest growth to measure biodiversity and climate change impacts.

In-house analysis

The sensor network is the plumbing. The policy decisions it informs are the palace. Who funds the next 1,000 nodes is the question.

Read our brief

The brief, by email

The AI moves that matter, filtered.

Approved stories, in-house analysis, and the operating consequence. One concise email. No generic roundup.