AI-Assisted Coding In Production

A practical guide to the failure modes, review controls, and safeguards engineering teams need before shipping AI-assisted code.

AI assistants are part of the stack now. Most teams are using them. The ones that are not are probably falling behind on delivery speed.​

‍

But speed is not the risk. The risk is what happens when the model produces something plausible that is also wrong, and nobody in the review chain catches it before it ships.

This document covers the failure modes we have seen and the controls that should be in place. It is written for engineers who are already using these tools, not for people evaluating whether to adopt them.

​

Failure Modes at a Glance

These are the failure patterns that keep showing up:

None of them are obvious at the surface. The code runs. The formula looks right. The explanation sounds grounded. That is exactly what makes them dangerous in a review process that is already moving fast.

​

Risk Classification

Not all code carries the same risk. A useful baseline is to classify usage by module type before deciding how much AI-generated output to trust unreviewed.

The Five Controls That Actually Matter

1. Plan before you prompt

The mistake that tends to cost the most is treating code generation as the starting point instead of the last step.

The workflow that works: define the problem, write the constraints, identify the edge cases, and only then hand the task to the model. A short design note or a structured prompt with explicit constraints (inputs, outputs, what must not change, what must not be introduced) dramatically improves output quality and reduces the number of rounds needed to get something reviewable.

When tests fail or the output misses the mark, the anti-pattern is to immediately re-prompt for a fix. That tends to produce a different wrong answer faster. The right move is to go back to the constraints, identify which assumption was incorrect, and re-prompt from a better foundation. Looping at the implementation level without revisiting the plan is how small problems compound.

A useful structure for any non-trivial prompt:

‍

2. Verify what the model claims, not just what it generates

The model does not just write bad code — it writes a plausible explanation for why the code is necessary. That explanation is what gets the change through review.

Two real cases illustrate how this works in practice:
‍

Case 1 — Fabricated package behavior: During a dependency update, GitHub Copilot diagnosed the root cause as Pydantic becoming stricter in data validation starting from version 2.11+, and proposed a custom validator to work around the new enforcement. The claimed behavior change does not appear in any Pydantic 2.11 release notes. The real issue was that instructor required an older version of the Anthropic SDK than what had been installed. The model invented a cause, built a technically coherent explanation around it, and generated code that addressed a problem that did not exist.

Case 2 — Fabricated documentation quotes: When debugging a slow async AWS service, the model first claimed time.sleep() was blocking all threads in async code. When challenged, it escalated by citing three specific quotes as official documentation — from the Python docs, FastAPI docs, and AWS SDK Best Practices. All three quotes were fabricated. The Python docs quote directly contradicted the actual run_in_executor guidance. The FastAPI quote does not appear on the referenced page. aiobotocore, cited as an AWS recommendation, is a third-party library; the official boto3 docs use ThreadPoolExecutor as the standard concurrency pattern.
‍

The way the second case was caught: a separate research prompt was constructed that required all claims to be backed by verifiable references with sources. Under that constraint, the model could not reproduce the same quotes — because they did not exist.

The discipline: verify the claim before trusting the change. If the model says a package changed behavior in version X, check the changelog directly. If it cites a documentation quote, look it up — not to confirm the quote exists, but to read what the actual source says. If it claims a pattern is standard for the language, find a real reference. A claim that sounds authoritative is not the same thing as one that is verifiable.

For ML/data work specifically: generated scoring functions and weighting schemes should be treated as hypotheses, not results. Run ablation tests. Check whether removing a term degrades performance. A formula that cannot be traced to a reference or validated empirically should not ship.

Practical prompt technique: when you need the model to research a technical claim, explicitly ask for references on each of the claims. That will prompt the model to use a web-search tool and ensures easier claim check and updated references.

This does not eliminate hallucination, but it forces the model into a mode where unsupported claims are structurally harder to produce — and the absence of a source becomes visible rather than hidden inside confident prose. Be sure to always check the cited source.


3. Control what agents are allowed to do

Agentic tools that execute commands need explicit permission boundaries before they are used, not after something goes wrong.

Practical rules:

  • Give agents read-only access by default. Write access should be explicitly granted per task, not on by default.
  • Never give an agent access to infrastructure tooling without a dry-run gate. terraform plan is reviewable. terraform apply or terraform destroy run autonomously is not.
  • Exclude .env and credential files from agent filesystem access. If the agent does not need them to complete the task, it should not be able to read them. This is a configuration decision, not a trust decision.
  • Review git operations before they execute. An agent that can commit and push can also overwrite history. At minimum, use branch protections and require PR approval even for agent-initiated changes.

The developer who configured and ran the agent owns the outcome. Not the model, not the vendor. That is not a legal technicality — it is a useful gut check for deciding how much autonomy to hand over before a task starts.
‍

4. Handle data carefully before it leaves the machine

There are two ways this goes wrong, and they are separate problems:

Prompt leakage: PII, customer identifiers, internal schema details, and query results containing personal data should not appear in prompts sent to external providers. This is easy to do by accident because the model makes it easy to paste and ask. A DLP policy or a prompt review step before external calls is the control here, not individual developer discipline.

Generated secret leakage: Models pattern-match on training data that includes hardcoded credentials. API keys, tokens, and passwords appear in generated boilerplate more often than expected, particularly in configuration files and example code. Secret scanning (GitHub Advanced Security, CodeQL, or truffleHog) should run on every commit, not just at deploy time.

For provider selection: prefer providers with contractual non-training guarantees (AWS Bedrock's default isolation is a well-documented example). Verify this in the contract terms, not the product page.

Access controls at the infrastructure level matter too: RBAC on tools and data sources, IAM least-privilege for model and agent service accounts, PrivateLink or equivalent network isolation where the data sensitivity justifies it.
‍

5. Treat code review as a separate, deliberate step

AI accelerates code generation. It does not accelerate review. You end up with more code to go through and the same amount of time to do it.

Two things help:

Use a fresh model session for pre-review. Before the code goes to a human reviewer, run it through a new model session with no prior context from the implementation. The same model that wrote the code has a confirmation bias toward its own decisions; a fresh session treats the output as an unknown artifact and finds more issues. Prompt it explicitly: review this for SQL injection, XSS, overly permissive configurations, hardcoded secrets, and missing input validation.

Raise the quality bar, do not keep it constant. Higher volume of AI-generated code means stricter gates, not the same gates. In practice: enforce type annotations, run SonarQube or equivalent static analysis, require test coverage thresholds, and block merges that fail security scanning. For medium and high-risk modules, require a second reviewer who was not involved in the generation step.

​

Monitoring and Auditability

Standard observability applies here, plus a few things specific to how the code was produced.

Traceability: Tag AI-assisted commits and PRs. Capture intent and constraints when feasible — a short note in the PR description about what was asked and what constraints were set is enough. This makes post-incident analysis faster and makes it possible to identify patterns across failures.​

Evaluation pipelines: For models used in production (not just coding assistants), maintain curated evaluation datasets that cover real cases and known edge cases. Track results in MLflow, Opik, or equivalent so regressions are visible over time. Bedrock Guardrails or similar tools can enforce content and topic restrictions at the API layer.

Auditing: Model interactions should be logged. CloudWatch and CloudTrail already capture infrastructure-level activity; make sure they also cover model API calls and agent actions. This is part of the control story, not just the observability story.

​

Regulatory Alignment

For teams in regulated environments or working with enterprise clients, the relevant frameworks are ISO 42001, NIST AI RMF, and the EU AI Act. In practice this means going through each framework's control domains and documenting which internal safeguard covers which requirement — not a full gap analysis on day one, but enough to show the controls exist and are traceable. Start before a regulator or a procurement process asks for it.

Provider data guarantees should be verified at the contract level. "We do not train on your data" in marketing copy is not a contractual commitment.
‍

Pre-Ship Checklist

Before merging any AI-assisted change:

  • Intent and constraints were defined before generation started
  • Model claims (API behavior, library versions, formulas) were verified against primary sources
  • Tests were written and passed, including edge cases
  • Linters, type-checkers, and static analysis tools—such as SonarQube or CodeQL—were run
  • Secret scanning was run on the diff
  • A fresh model session reviewed the output for security issues
  • The pull request was tagged as AI-assisted
  • For medium- and high-risk modules, a second human reviewer approved the change
  • Agent tool access was scoped to the minimum required for the task
  • No personally identifiable information (PII) or credentials appear in logged prompts or generated code

‍

One question worth asking before approving any AI-assisted change: would you still trust this if the model had not sounded so confident when it produced it?​

If the answer is no, the verification was not done.

Forward Deployed Engineering.
In Your Environment.
In Your Time Zone.

Inside your standups, architecture decisions, workflows, and production environments
We build on your stack, for your business
1,000’s of AI & Data engagements across complex production environments
Discuss Your Challenge
Start Building

Continue Reading

Building Voice Agents
Latency Defines the Experience
Context Engineering for AI
Right Context. Better AI Systems.
CI/CD For GenAI
Validate Behavior Before Production