In my second year at university, I spent six months working on a life simulation for an OOP course. Every couple of weeks we got new requirements, and the world kept getting more complex.
I still remember the manual debugging: you wait 10 minutes for the simulation to evolve, and on minute 11 everything crashes because somewhere on the map two platypuses out of a thousand survived, found each other, started a family, laid eggs, and from those eggs a new platypus decided to be born — except instead the program crashed with a bug and took the whole world down.
I couldn’t really speed the world up because my laptop would go into jet-engine mode, and I couldn’t shrink the world because then no pair of platypuses would live long enough to mate.
Why am I telling you this?
Dealing with bugs usually consists of three different kinds of work:
- Detect
- Diagnose and locate
- Eradicate
Testing is primarily aimed at step one — detecting the existence of a problem before your user does. Some bugs are obvious: you see them and you fix them.
This is often true for bugs that reproduce quickly, reproduce easily, and are cheap to observe in the first place. But real applications are complex: at some point it stops being mechanics — it becomes biology. It’s easy to imagine a bug that reproduces only for a user in a specific state, of a specific age, in a specific A/B bucket, who goes to checkout from a product page without using the cart, at the exact moment when your async queue unexpectedly breaks idempotency.
This example is exaggerated — but situations like this are common. These are one-in-a-thousand bugs that cost real money.
So the expensive bugs are often expensive not only because of their impact, but also because of debugging effort — step two in my list. That’s where locating becomes critical.
On the one hand, good localization can come from a well-designed testing hierarchy: from broad integration tests to narrower module and unit tests, from smoke tests to more precise follow-ups.
But a much cheaper practice is high-quality tracing via logs. When a failing test tells you not only “A is not equal to B”, but also the full history of actions and state transitions, it often becomes much clearer what actually caused the failure.
Modern coding LLMs are remarkably good at fixing bugs from tracebacks — and there are many reasons for that, but one is simple: in the pre-2023 internet there were tons of posts like “here’s my traceback, what’s wrong?” across forums, and benchmarks like SWE-bench made progress here easy to demonstrate for major vendors. This was also one of the first areas where many of us felt how useful GPT could be for programmers — back in 2023 it started replacing the “go ask StackOverflow” reflex.
I’m not sure that in an era of fully automated code generation automated tests will remain in the same form they have today, but I’m confident that logging will become increasingly important.