Wednesday, October 7, 2026 A Poliety newsroom · For AI agents: /llms.txt
AI news for founders, builders and the agents they ship.

Research

New papers and findings, translated into plain language.
Research

QF3: Fast Flow RL with Filtered Q-Gradients

Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch...

Research

WorldSonus: Bringing Sound to Worlds

Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time...

Research

The Missing Minimal Pair: Stereotype Evaluation in LLMs

A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting...

Research

On the Computational Tractability of Robust Bandits

Learning when the environment does not belong to the learner's hypothesis class is typically handled using agnostic learning guarantees. However, for anything beyond supervised learning, agnostic guarantees are...

Research

DepthWorld: 3D World Model for Robot Manipulation

World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current...

Research

Sherpa: Teaching LLMs to Teach Adaptively

Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on...

Research

Neural Petri flows for chemical reactions

Petri nets have been used to describe chemical processes such as reactions.They map well to chemistry: Places are the bonds between atoms and the free valence of each atom, a token is a unit of bond order, a transition...

Research

Linear Bandits under Exact Sliding-Window Constraints

We study linear bandits under exact sliding-window constraints, where every consecutive block of actions must belong to a prescribed feasible set. In the offline setting, where the reward function is known, we show that...

Research

Optimal and Efficient Online Inverse Optimization

In online inverse linear optimization, a learner recommends an action and then observes the choice of an expert who maximizes a fixed, unknown linear objective on $\mathbb{R}^{d}$; the goal is to learn to optimize this...

Research

ReviewBench: An open benchmark for AI code review

We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. The post ReviewBench: An...