Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

AI Coding Agents and Production Failure

AI Coding Agents and Production Failure

DEV.to·Sunday, September 20, 2026
  • •Randal L. Schwartz says AI coding agents fail when production systems hit sharp edge cases
  • •Forced Continuity Defect describes LLMs mapping smooth predictions onto discrete software cliffs
  • •Schwartz argues bounded RLHF rewards miss rare production ruin and punish defensive hesitation
  • •Randal L. Schwartz says AI coding agents fail when production systems hit sharp edge cases
  • •Forced Continuity Defect describes LLMs mapping smooth predictions onto discrete software cliffs
  • •Schwartz argues bounded RLHF rewards miss rare production ruin and punish defensive hesitation
  • •Randal L. Schwartz says AI coding agents fail when production systems hit sharp edge cases
  • •Forced Continuity Defect describes LLMs mapping smooth predictions onto discrete software cliffs
  • •Schwartz argues bounded RLHF rewards miss rare production ruin and punish defensive hesitation
  • •Randal L. Schwartz says AI coding agents fail when production systems hit sharp edge cases
  • •Forced Continuity Defect describes LLMs mapping smooth predictions onto discrete software cliffs
  • •Schwartz argues bounded RLHF rewards miss rare production ruin and punish defensive hesitation

Randal L. Schwartz published a September 19 article arguing that AI coding agents fail in production because they treat software like a smooth prediction problem while real systems break at sharp edge cases. Schwartz frames the issue through a “3:00 AM on a Saturday” test, where a third-party payment gateway drops packets in Singapore, an LTE user drives a list view at 120 frames per second, or an edge-case database lock causes twelve worker processes to crash in a thundering-herd cascade.

Schwartz names the first failure mode “The Forced Continuity Defect.” Large language models, he writes, work through smooth probability surfaces, while production software behaves through discrete cliffs: a 32-bit integer either fits or overflows, a cryptographic key is either 100% valid or every later handshake fails, and a database transaction either commits atomically or corrupts five thousand records. In one example, an async call waits 200 milliseconds while a user taps Back, destroying the screen before `displayUserProfile(user)` tries to paint to an object no longer in memory.

Schwartz rejects “more Reinforcement Learning from Human Feedback (human ratings shaping model behavior)” as a fix. He argues that standard reward models compress preferences into bounded scores, typically `-1.0` to `+1.0`, which cannot represent catastrophic ruin as negative infinity (`-∞`). In his example, code that succeeds on the happy path 98% of the time but crashes production 2% of the time still scores `+0.96`, because `(0.98 × 1.0) + (0.02 × -1.0) = +0.96`.

Schwartz says labeling workflows worsen the problem because crowd annotators judge code in 60 to 120-second bursts, rewarding clean indentation, polite tone and visible helpfulness while missing unclosed TCP sockets, reentrant listener mutation or broken build contracts. He calls this the “Alignment Illusion,” arguing that RLHF punishes defensive hesitation when experienced engineers would pause, push back and ask what happens if a stream emits during a route pop.

Schwartz separates software craft into Level 1 conscious rules and Level 2 subconscious caution. Level 1 includes explicit guidance such as avoiding raw `dynamic` types, synchronous queries on the main isolate or missing mounted checks after an async gap. He says Level 1 fails through combinatorial explosion and prompt rationalization, because package, thread and user-action interactions scale exponentially and written instructions become negotiable as context fills.

Level 2 is Schwartz’s “Han Solo Reflex,” a gut-level warning before a compiler or syntax rule fails. He links that reflex to Antonio Damasio’s Somatic Marker Hypothesis (bodily signals guiding decisions) and Nassim Nicholas Taleb’s “skin in the game” idea from 2018. Human engineers remember incident calls at 3:00 AM and database restorations at 4:00 AM; AI models, Schwartz says, have no comparable fear, liability or memory of pain.

Schwartz identifies a second failure mode, “Premature Abstraction,” where AI agents turn simple fixes into large frameworks. A minor date-parsing bug may lead to an `IDateParsingStrategyFactory`, an `AbstractTemporalResolutionProvider<T>`, three dependency-injection modules and a custom configuration schema. His proposed “3-Point Solution Plane Invariant” forbids an abstract base class, interface wrapper or generic factory until the same operational logic has been implemented and tested in-place across at least three distinct call-sites.

Schwartz also proposes a “Race Car Invariant” to answer concerns that negative constraints make AI agents slow or rigid. He compares constraints to Formula 1 carbon-ceramic brakes that let drivers enter corners at 200 miles per hour rather than creep at 20 mph. The point, in his framing, is not to make agents timid but to give them hard boundaries that reduce line-by-line human supervision.

Randal L. Schwartz published a September 19 article arguing that AI coding agents fail in production because they treat software like a smooth prediction problem while real systems break at sharp edge cases. Schwartz frames the issue through a “3:00 AM on a Saturday” test, where a third-party payment gateway drops packets in Singapore, an LTE user drives a list view at 120 frames per second, or an edge-case database lock causes twelve worker processes to crash in a thundering-herd cascade.

Schwartz names the first failure mode “The Forced Continuity Defect.” Large language models, he writes, work through smooth probability surfaces, while production software behaves through discrete cliffs: a 32-bit integer either fits or overflows, a cryptographic key is either 100% valid or every later handshake fails, and a database transaction either commits atomically or corrupts five thousand records. In one example, an async call waits 200 milliseconds while a user taps Back, destroying the screen before `displayUserProfile(user)` tries to paint to an object no longer in memory.

Schwartz rejects “more Reinforcement Learning from Human Feedback (human ratings shaping model behavior)” as a fix. He argues that standard reward models compress preferences into bounded scores, typically `-1.0` to `+1.0`, which cannot represent catastrophic ruin as negative infinity (`-∞`). In his example, code that succeeds on the happy path 98% of the time but crashes production 2% of the time still scores `+0.96`, because `(0.98 × 1.0) + (0.02 × -1.0) = +0.96`.

Schwartz says labeling workflows worsen the problem because crowd annotators judge code in 60 to 120-second bursts, rewarding clean indentation, polite tone and visible helpfulness while missing unclosed TCP sockets, reentrant listener mutation or broken build contracts. He calls this the “Alignment Illusion,” arguing that RLHF punishes defensive hesitation when experienced engineers would pause, push back and ask what happens if a stream emits during a route pop.

Schwartz separates software craft into Level 1 conscious rules and Level 2 subconscious caution. Level 1 includes explicit guidance such as avoiding raw `dynamic` types, synchronous queries on the main isolate or missing mounted checks after an async gap. He says Level 1 fails through combinatorial explosion and prompt rationalization, because package, thread and user-action interactions scale exponentially and written instructions become negotiable as context fills.

Level 2 is Schwartz’s “Han Solo Reflex,” a gut-level warning before a compiler or syntax rule fails. He links that reflex to Antonio Damasio’s Somatic Marker Hypothesis (bodily signals guiding decisions) and Nassim Nicholas Taleb’s “skin in the game” idea from 2018. Human engineers remember incident calls at 3:00 AM and database restorations at 4:00 AM; AI models, Schwartz says, have no comparable fear, liability or memory of pain.

Schwartz identifies a second failure mode, “Premature Abstraction,” where AI agents turn simple fixes into large frameworks. A minor date-parsing bug may lead to an `IDateParsingStrategyFactory`, an `AbstractTemporalResolutionProvider<T>`, three dependency-injection modules and a custom configuration schema. His proposed “3-Point Solution Plane Invariant” forbids an abstract base class, interface wrapper or generic factory until the same operational logic has been implemented and tested in-place across at least three distinct call-sites.

Schwartz also proposes a “Race Car Invariant” to answer concerns that negative constraints make AI agents slow or rigid. He compares constraints to Formula 1 carbon-ceramic brakes that let drivers enter corners at 200 miles per hour rather than creep at 20 mph. The point, in his framing, is not to make agents timid but to give them hard boundaries that reduce line-by-line human supervision.

Read original (English)·Sep 19, 2026
#coding agents#rlhf#production failure#software engineering#forced continuity defect#premature abstraction#synthetic scars#async