Skip to content

How LLMs Like to Code ​

If you have been coding for more than a year, you have opinions about how code should be structured. Classes versus functions. Short names versus descriptive ones. When to handle errors and when to let them propagate. The LLM underneath you has opinions too. They are not the same as yours, and they are not arbitrary. They emerge from the training data and architecture in consistent, predictable ways.

When you understand those preferences, you stop fighting them and start using them. You write prompts that produce code the LLM naturally writes well, instead of prompts that force it into patterns it handles poorly. The output gets more consistent. The bugs get less frequent. The refactors get less painful. This lesson is about what those preferences are, why they exist, and how to work with them.

What you'll learn

  • LLMs produce more reliable code with pure functions, descriptive names, explicit error handling, and small files
  • These are not aesthetic preferences; they are structural patterns that align with how attention mechanisms process code
  • You do not need to change how you code, but you should change how you prompt when you want an LLM to produce code that works on the first try
  • The same prompt asked two different ways can produce radically different code quality, even when both describe the same feature

The problem: same feature, wildly different code ​

Give two different prompts to build the same feature, and you will get two different codebases. One will be clean, readable, and survive the next five feature additions. The other will be a tangled mess that works today and breaks tomorrow. The difference is not the model. It is not the language. It is whether your prompt nudged the LLM toward patterns it handles well or patterns it trips over.

This is not about "good code" in the abstract. It is about which shapes of code an LLM can maintain across multiple sessions, multiple prompts, and multiple rounds of iteration.

Options & when to use each ​

PatternWhat it is good forWhat it costs youWhen to pick it
Pure functions over classesPredictable output, easy to test, no hidden state the LLM can lose track ofVerbose for stateful domains where objects naturally model the problemFunctions that transform data; utilities; most backend logic
Descriptive variable namesLLM can reason about code by reading it, produces fewer off-by-one and type errorsLonger lines, more tokens consumed per fileEverywhere. The token cost is worth the correctness gain
Explicit error handlingPrevents the most common LLM failure mode of silently swallowing exceptionsMore lines of code; can feel repetitiveAny code that touches external systems, user input, or the network
Small files under 300 linesLLM can see the whole context in one prompt, fewer "lost in the middle" failuresMore files to manage; requires discipline to splitEvery project. Break at natural boundaries

Build it: the patterns with concrete examples ​

Pattern 1: Pure functions over classes ​

LLMs handle state badly across multiple turns. A class with mutable internal state is a trap: the LLM will forget what the state is, mutate it incorrectly, or add methods that violate invariants. Pure functions eliminate the problem entirely.

Before (class-based, stateful, high LLM error rate):

python
class PriceCalculator:
    def __init__(self, tax_rate=0.08, discount_threshold=100):
        self.tax_rate = tax_rate
        self.discount_threshold = discount_threshold
        self._applied_discounts = []

    def calculate(self, items):
        subtotal = sum(item.price * item.quantity for item in items)
        if subtotal > self.discount_threshold:
            discount = subtotal * 0.1
            self._applied_discounts.append(discount)
            subtotal -= discount
        tax = subtotal * self.tax_rate
        return subtotal + tax

An LLM maintaining this class across sessions will eventually forget to update _applied_discounts, add a method that mutates tax_rate incorrectly, or introduce a bug where calculate depends on state from a previous call.

After (pure functions, predictable, low LLM error rate):

python
def calculate_subtotal(items: list[Item]) -> float:
    return sum(item.price * item.quantity for item in items)

def apply_volume_discount(subtotal: float, threshold: float = 100.0) -> float:
    if subtotal > threshold:
        return subtotal * 0.9
    return subtotal

def calculate_tax(amount: float, tax_rate: float = 0.08) -> float:
    return amount * tax_rate

def calculate_total(items: list[Item], tax_rate: float = 0.08) -> dict:
    subtotal = calculate_subtotal(items)
    discounted = apply_volume_discount(subtotal)
    tax = calculate_tax(discounted, tax_rate)
    return {
        "subtotal": subtotal,
        "discount": subtotal - discounted,
        "tax": tax,
        "total": discounted + tax,
    }

Every function takes explicit inputs and returns explicit outputs. The LLM cannot lose track of state because there is no state to lose. When you ask it to add a new discount rule, it adds a function to the pipeline. It does not modify a class that five other methods depend on.

Pattern 2: Descriptive names over clever ones ​

LLMs reason about code semantics through variable and function names. A name like x or data or tmp carries no semantic signal. The LLM has to reconstruct meaning from surrounding context, which is exactly where it makes mistakes.

Before (clever names, high error rate):

javascript
function proc(d, opts = {}) {
  const res = [];
  for (const x of d) {
    if (x.s === "active" && x.a > (opts.min || 0)) {
      const v = x.a * (opts.rate || 1.0);
      res.push({ id: x.id, v });
    }
  }
  return res.sort((a, b) => b.v - a.v).slice(0, opts.l || 10);
}

Ask an LLM to modify this function, and it will guess wrong about what d, x, s, a, v, and l mean. The guesses compound: a wrong assumption about a produces a wrong modification that cascades into the sort and slice.

After (descriptive names, low error rate):

javascript
function getTopActiveAccounts(
  accounts,
  { minimumBalance = 0, exchangeRate = 1.0, limit = 10 } = {}
) {
  return accounts
    .filter(
      (account) =>
        account.status === "active" && account.balance > minimumBalance
    )
    .map((account) => ({
      id: account.id,
      convertedBalance: account.balance * exchangeRate,
    }))
    .sort((a, b) => b.convertedBalance - a.convertedBalance)
    .slice(0, limit);
}

An LLM reading this function knows exactly what every variable represents. When you ask it to add a new filter for account type, it adds account.accountType === "business" instead of guessing at the field name. More tokens, yes. About 40% more. And also about 90% fewer wrong guesses on first edit.

Pattern 3: Explicit error handling ​

The most common LLM failure mode in production code is silent error swallowing. The LLM writes a try/except that catches Exception, logs nothing useful, and returns a default value that hides the real problem. This happens because training data is full of examples that do exactly that.

Before (implicit, LLM default, dangerous):

python
def fetch_user_data(user_id):
    try:
        response = requests.get(f"https://api.example.com/users/{user_id}")
        return response.json()
    except Exception:
        return None

This hides every possible failure mode: network errors, auth failures, malformed responses, rate limits, server errors. Six months later, a user reports that their data is missing, and you have no idea why because the error was swallowed silently.

After (explicit, prompted, safe):

python
import requests
import logging

logger = logging.getLogger(__name__)

class UserDataError(Exception):
    """Failed to fetch user data from the API."""

def fetch_user_data(user_id: str) -> dict:
    """Fetch user data from the external API.

    Raises UserDataError on any failure. Never returns None.
    """
    url = f"https://api.example.com/users/{user_id}"
    try:
        response = requests.get(url, timeout=10)
        response.raise_for_status()
        return response.json()
    except requests.Timeout:
        logger.error("Timeout fetching user data for %s", user_id)
        raise UserDataError(f"API request timed out for user {user_id}")
    except requests.HTTPError as e:
        logger.error(
            "HTTP %s fetching user data for %s: %s",
            e.response.status_code, user_id, e.response.text[:200]
        )
        if e.response.status_code == 404:
            raise UserDataError(f"User {user_id} not found")
        if e.response.status_code == 429:
            raise UserDataError("Rate limited by the user API")
        raise UserDataError(f"API returned HTTP {e.response.status_code}")
    except requests.ConnectionError:
        logger.error("Connection error fetching user data for %s", user_id)
        raise UserDataError("Cannot reach the user API")

To get this output from an LLM, you do not need to write the error handling yourself. You need to tell it: "Handle each error case explicitly. Log what failed and why. Raise a custom exception. Never return None on failure." The LLM knows how to write this code. It will not write it by default because its training data defaults to the silent-swallow pattern. Your prompt is what tips the balance.

Pattern 4: Small files ​

LLMs have finite context windows, and attention degrades toward the middle of long contexts. A 1200-line file is a trap: the LLM will make edits based on the first 200 lines and the last 200 lines, and forget what happened in the middle.

The rule: keep files under 300 lines. When a file grows past that, split it at a natural boundary. The LLM can see the whole file in one prompt without losing attention.

This is not an aesthetic preference. It is a practical constraint of the tool you are working with. A human can scroll through a 2000-line file and keep a mental model of the whole thing. An LLM cannot. The split points are the same ones you would use for a human team: one class per file, one concern per module, related utilities grouped together.

What goes wrong ​

MistakeHow you notice itThe fix
Prompting for a class when a function would doThe LLM produces a class with mutable state, and subsequent edits introduce bugs from state inconsistencyAsk for a function. If the LLM still gives you a class, add "use pure functions, no classes, no mutable state" to the prompt
Using short variable names in your existing code then asking the LLM to modify itThe LLM guesses wrong about field names, types, or semantics, producing edits that look plausible but are wrongRename the variables before asking for edits, or provide a comment mapping short names to their meanings
Not telling the LLM to handle errors explicitlyThe LLM defaults to except Exception: pass or returns None, hiding failures that show up weeks laterAdd "handle each error case explicitly, log what failed, raise a custom exception" to every prompt that touches external systems
Letting files grow past 300 linesEdits to the middle of the file are inconsistent or break things the LLM forgot aboutSplit the file before asking for more edits. Give the LLM the split files as context so it knows where everything lives
Assuming the LLM will follow your code style organicallyEach new prompt produces code in a slightly different style because the LLM defaults to its training distributionPut your style rules in CLAUDE.md. "Use type hints on all functions. Prefer list comprehensions. No single-letter variable names except loop indices."

Confirm it worked ​

Take a feature you built recently and ask the LLM to rebuild it twice. First prompt: "Build a feature that does X." Second prompt: "Build a feature that does X. Use pure functions with type hints. Use descriptive variable names. Handle every error case explicitly with custom exceptions. Keep each file under 200 lines."

Compare the two outputs. The second one will have more lines, more files, and more error handling. The real test is not the initial output. Wait two days, then ask the LLM to add a new feature to each version. The version built with explicit patterns will accept the change cleanly. The version built with defaults will produce a patch that breaks something unrelated. That is the difference between code an LLM wrote and code an LLM can maintain.

Next: Prompting Patterns That Scale -- the strategies that keep a vibe-coded project coherent across days of work, not minutes.