reviewing code with a machine

you are absolutely right, the buffer is definitely overflowing

I spend a lot of my working day reading other people’s code. Some hear ethical hacker and think HACKERMAN, but usually it means staring at a function, figuring out what it does and whether I can make it do something it shouldn’t. In a busy month I review 2 to 4 custom codebases for security vulns, though the more complex ones can eat up the entire month.

Getting into a codebase used to take a while. I would clone the thing and start building a mental map, chasing the odd choices that even the original developers can no longer explain and following functions until I got why they existed. I hear a lot of people say they would rather write code than review it, but I love this part.

LLMs have changed how people write code and how it gets reviewed, which means that initial period I love is getting shorter. Plenty of software engineers have written about the sense of loss that comes with their field changing, and I feel some of that too, even while enjoying the fact that I can now spend an afternoon building a rough mental map that used to take me days. It is a strange thing to miss.

The model reads faster than I do (I am dyslexic after all), it never gets bored (I have ADHD after all), and it will happily trace a call path across forty files while I go make coffee. I do have to be careful how I ask, because it can take things very literally, which, as someone with autism, I feel I should probably be more patient with.

the fast part

It is very good at the monotonous work that used to cost me entire days: looking for every place a function gets called, tracing what touches a piece of user input, checking which of forty near-identical handlers actually reach the database.

A recent one was a typical app where each route handler had to call the access-control check itself, which meant every route was one forgotten line away from being public. Verifying that by hand is a day of reading the same function shape on repeat. Now it is asking the model to walk every route and flag the ones missing the check, and then making that coffee. It came back with a short list of routes, most of which were public on purpose. One was not supposed to be.

confidently wrong

The counterargument I hear the most is that these models (not picking favourites now, that is a discussion for another day) will hand you a wrong view of the code. And yes, sometimes it misses the point of a module entirely, and it is confidently wrong.

The tricky part is that it does not tell you when it is guessing. A handler it traced properly and a handler it made up on the spot get explained the same way, clean, plausible, well structured, same confident tone.

My favourite example so far was a memory issue it found in a recent review. I asked it to verify itself and it came back with no, that was a mistake, there is no issue there. Asked it to double check a final time just to make sure: “You are absolutely right! I see it clearly now, yes, the buffer is definitely overflowing here.” Same code, full confidence every round. It cannot even agree with itself.

The overflow was real, and I ended up using AI to help build and debug a PoC for it. Giving it that concrete goal got me further than asking it for another opinion on the same code, because now we had something to run and check. Still using the thing, after all that.

Some wrong answers are easy to spot because they do not make sense when I read them, so I go look for myself. The fluent ones are the sneaky ones, because they read exactly like correct ones. In an assessment that can cost you a finding if you wave off a suspicious path because a made-up summary said it was fine.

So I do hear the counterargument, it really is wrong a good chunk of the time. But so am I. Half of my early theories about a codebase die as I read the rest of it, which has always been part of the job.

The summaries make that harder in one particular way: it is tempting to move on as if I read the code myself when I only read a summary of it. For anything that matters I go read the actual code, even when that means chasing something the model should never have sent me towards in the first place.

Some of the wrong answers are still useful, especially when it read the right files and traced a real path before taking a wrong turn somewhere along the way. I can work with that, and sometimes there is an original idea hiding in it or a line of thinking worth following myself. A bit of interactive rubber ducking.

And the slow way did not go anywhere. I still build intuition by following a function by hand, getting confused and backtracking, sometimes carrying three half-wrong theories about a module until one of them survives contact with the code. It is exactly the kind of rabbit hole my brain loves to disappear into (the ADHD again). I just get a head start now, which leaves more time to look at other things.

still reading code

Mostly it has shifted the job, with less of my time going into the slow first read, the part I said I love, and more going into deciding what to actually verify. Turns out I love that too.

I wrote in the first post here that I like tools but that the thinking has to stay mine. That still holds, even if I have to work harder and faster these days to keep it that way.