Run one of Anthropic or OpenAI's coding agents for a couple of hours on a real repo and what you get back is a chat transcript, which tells you almost nothing about what it actually changed on disk.
Both labs keep pushing how long these agents can run on their own, and neither ships a matching way to reconstruct what one did after the fact. You can watch the whole run happen live and still not be able to say, a day later, which one of its edits broke prod.
Both labs keep pushing how long these agents can run on their own, and neither ships a matching way to reconstruct what one did after the fact. You can watch the whole run happen live and still not be able to say, a day later, which one of its edits broke prod.
