The views expressed here are my own and do not represent those of OpenAI. The examples are illustrative and do not describe OpenAI’s internal practice
The views expressed here are my own and do not represent those of OpenAI. The examples are illustrative and do not describe OpenAI’s internal practice

The easiest way to make bad work today is to confuse activity with care. Sometimes that produces obvious slop. Sometimes it produces something stranger: sloppy perfectionism.

An agent can write a feature in minutes. Another can review it. A third can find something else to improve. Before long, there is code, feedback, revisions, and a perfectly respectable-looking trail of work. What there may not be is anybody who stopped to ask whether any of it matters.

We have become extraordinarily good at producing things. We are still learning how to judge them. Knowing when something is good enough has always been difficult. With AI, it has become harder, easier to avoid, and more valuable than ever.

Two ways not to care

Sloppy work is usually easy to recognize. The product that gets the basics wrong. The deck with three different fonts. The feature whose rough edges make it clear that nobody tried using it before it shipped.

These things send a message. Somebody wanted the work to be done more than they wanted it to be useful.

Perfectionism sends a different message, but it can come from the same place.

If you spend weeks polishing a feature before finding out whether anybody needs it, you are not necessarily being careful. You may simply be ignoring a different set of consequences: the customer still waiting, the feedback you have not gathered, the other problems your team could have solved.

One kind of carelessness fails to take the work seriously. The other fails to take its purpose seriously.

The most valuable work does not come from applying the highest possible standard everywhere. Nor does it come from splitting the difference and declaring the middle good enough. It comes from understanding what this particular piece of work is supposed to achieve, and giving it the care that purpose demands.

That is the slop frontier.

Article content
The slop frontier: the right quality bar depends on context, not a fixed midpoint.

Sometimes that frontier sits close to perfection. A sales deck for a product that promises exceptional quality cannot afford sloppy typography or an incoherent visual language. The details are part of the argument. They tell the reader what kind of people built the product and whether those people can be trusted.

Sometimes the bar belongs somewhere else entirely. A prototype used once in a research study needs to answer the research question. If it does that without misleading participants, time spent polishing the rest may be time taken away from learning.

Same principle. Completely different standard.

Good enough has a reputation problem

“Good enough” often sounds like something you say when you have run out of time or stopped caring. It has the faint air of an excuse, as if the person saying it would obviously have done better if only they had tried harder.

I think that gets it backwards.

Knowing when something is good enough is an acquired skill. It requires understanding the purpose of the work, the people affected by it, the risks involved, and the cost of getting the judgment wrong. It means knowing which details change the outcome and which ones merely invite another round of tinkering.

Stopping too early is easy. Continuing indefinitely is easy in its own way, too. Both let you avoid the uncomfortable decision about where the bar actually belongs.

The hard part is making that decision and being able to explain it.

Products make this especially clear. The quality bar for the first experience, the core workflow, or a moment that determines whether someone trusts the product may need to be extremely high. A new feature that is still searching for its shape needs something different. It should get the basics right, but it probably does not need weeks of refinement before the first person sees it.

A pixel out of place is not automatically a problem. A dozen small details that add up to “nobody thought about me” usually are. The difference is not something you can settle with a checklist. It takes context, experience, and a sense for how the work will land with the person on the other side.

That is what good enough actually means: a judgment about the right amount of care, in the right place, at the right time.

AI makes the easy mistake easier

For a long time, poor-quality work at least took some effort to produce. A half-baked feature still had to be written. A bloated document still required someone to sit down and fill the pages. The cost of making things imposed a natural, if imperfect, limit.

That limit has shifted.

You can now one-shot a feature, generate a presentation, or ask an agent to draft a response in moments. The output often looks complete before anybody has decided whether it is correct, useful, or appropriate.

The pressure around the work has changed, too. When producing something becomes easier, the expectation becomes that we should produce more of it. More features, more reviews, more responses; less time to think about each one.

Under those conditions, the quality bar is more likely to fall than to rise.

Niklas Gruhn captured one version of the problem in Don’t be a meat proxy. A meat proxy sits between an AI system and another person, passing the output along without reading it, understanding it, or adding any judgment. The recipient inherits the work of figuring out whether the response makes sense.

The same thing happens with code. Someone hands a ticket to an agent, forwards the resulting change for review, and feeds every comment back into the agent without ever developing a view of the work themselves. The feature may eventually ship. But the thinking did not disappear; it was offloaded to the reviewers.

The problem is not using agents. The problem is giving up the part of the work that makes your involvement valuable.

You do not want to be a meat proxy. But avoiding that trap does not mean obsessing over every detail an agent can find. There is another, stranger failure mode waiting in that direction.

Sloppy perfectionism

Ask an agent to review a piece of work, and it will usually find something.

A sentence could be more concise. A variable could have a different name. An unlikely edge case might be handled more elegantly. A paragraph could be reorganized. Run the review again, and there will be more. Run a second agent, and it will find different things.

Some of those observations will matter, but many will not.

The failure begins when nobody makes that distinction. An agent produces a list of issues. A human forwards the list. Somebody else spends an afternoon fixing it. Another agent reviews the result and discovers another set of issues. Everybody involved can point to the process and say that the quality bar is being upheld.

But what bar?

If nobody has examined whether the findings are relevant, whether the proposed fixes improve the outcome, or whether the work was already good enough for its purpose, there is no meaningful standard being enforced. There is only an endless supply of things an agent was capable of noticing.

That is sloppy perfectionism.

It has the outward appearance of rigor: more reviews, more comments, more revisions, more insistence on quality. Underneath, it is the same abdication of judgment as shipping the first thing an agent produced.

The meat proxy forwards unexamined output as work. The sloppy perfectionist forwards unexamined output as standards.

Both make somebody else do the thinking.

That is what makes this failure mode so difficult to spot. Slop usually announces itself. Sloppy perfectionism arrives dressed as conscientiousness. It can make an entire team slower whilst giving everyone the comforting impression that they are being thorough.

A review is not valuable because it found something. A review is valuable when it found something that matters.

Judgment is the work

AI does not remove the tradeoff between sloppiness and perfectionism. It makes both ends cheaper to reach.

You can ship something nobody has properly considered. You can also spend days chasing improvements that don't matter. The first leaves customers or colleagues to deal with unfinished thinking, the second buries them in unnecessary work.

The answer cannot be to stop using agents, and it cannot be to blindly accept every suggestion they produce. It has to be a better understanding of what the work is for.

What must be true before this can help somebody? Which details will change the outcome? Which rough edges are harmless, and which ones quietly signal that nobody cared? What are we giving up if we spend another day on this? And if a review finds a problem, is it actually a problem here?

These questions have always mattered. They matter more when execution is cheap, output is abundant, and there is always another agent willing to suggest one more improvement.

The frontier moves with context. With the people involved. With the risks. With what you are trying to learn, build, sell, or change. Knowing where it sits is the work.

Good enough is not what happens when you stop caring. It is what happens when you care enough to know what matters.