That depends on how carefully youre checking and what youre checking for. I’ve seen debugging a single problem with a project take nearly as long as the whole rest of the project, so I don’t find that claim broadly true. Again, reality supports my observation because there are no companies turbocharging their productivity this way, or the people selling the AIs would make absolutely sure everyone sees that company doing it and its rising revenue and value to match.
Because the LLMs are error-prone, and the the types of errors they may make are not limited to the types a human is likely to make, you can’t do the standard supervisor “glance, nod, say it looks good” routine you might do with an apprentice’s work on a basic task, but LLMs can make any kind of mistake anywhere, so I have to use a fine-toothed comb or risk rookie mistakes on even rookie work submitted with my name on them.
Like I said before and still contend, AIs today are great for tasks where mistakes are not events and high accuracy is not critical to success. I think for anything not like that, they’re bad and are constrained by their very design to stay that way. They are no different than any other tool: good at a specific type of task and bad at everything else. _Un_like any other tool, though, widespread attempts are being made to use it for a broad variety of things that it is bad for in one or more ways.
ETA: In this post, “bad” can also include inefficient, overkill, or wasteful, like killing a bug with a bomb.