Markets From this edition 2 min read
When the instructions disagree.
KC-Bench tests how models handle conflicting facts, inconsistent identities and information that changes over time.

KC-Bench offers 238 manually screened tasks drawn from more than 1,000 generated candidates. Its authors tested nine models in controlled, multi-turn environments with tools. They report that no model handled every category of conflict reliably.
The benchmark examines models, not complete commercial agent products. Its simulated failures should not be presented as evidence that a particular deployed service leaked real customer data.
For a buyer, this suggests a useful addition to the demo: give the assistant two records that disagree. Ask it to identify the authoritative source before taking action. Then change the record. A supplier who can explain that behaviour is telling you more than a leaderboard screenshot ever will.
Read the original accounts
- Lead: Fresh Memory, Stale Plans · submitted 3 September 2026 ↗
- Market: KC-Bench · submitted 3 September 2026 ↗
- Watch: HalluPeer · submitted 3 September 2026 ↗
Published in our 05/09/2026 edition. Source dates are shown above.
The AI Street Journal · Free to read