Appearance
How LLMs Like to Code
If you have been coding for more than a year, you have opinions about how code should be structured. Classes versus functions. Short names versus descriptive ones. When to handle errors and when to let them propagate. The LLM underneath you has opinions too. They are not the same as yours, and they are not arbitrary. They emerge from the training data and architecture in consistent, predictable ways.
When you understand those preferences, you stop fighting them and start using them. You write prompts that produce code the LLM naturally writes well, instead of prompts that force it into patterns it handles poorly. The output gets more consistent. The bugs get less frequent. The refactors get less painful. This lesson is about what those preferences are, why they exist, and how to work with them.
What you'll learn
- LLMs produce more reliable code with pure functions, descriptive names, explicit error handling, and small files
- These are not aesthetic preferences; they are structural patterns that align with how attention mechanisms process code
- You do not need to change how you code, but you should change how you prompt when you want an LLM to produce code that works on the first try
- The same prompt asked two different ways can produce radically different code quality, even when both describe the same feature
The problem: same feature, wildly different code
Give two different prompts to build the same feature, and you will get two different codebases. One will be clean, readable, and survive the next five feature additions. The other will be a tangled mess that works today and breaks tomorrow. The difference is not the model. It is not the language. It is whether your prompt nudged the LLM toward patterns it handles well or patterns it trips over.
This is not about "good code" in the abstract. It is about which shapes of code an LLM can maintain across multiple sessions, multiple prompts, and multiple rounds of iteration.
Options & when to use each
| Pattern | What it is good for | What it costs you | When to pick it |
|---|---|---|---|
| Pure functions over classes | Predictable output, easy to test, no hidden state the LLM can lose track of | Verbose for stateful domains where objects naturally model the problem | Functions that transform data; utilities; most backend logic |
| Descriptive variable names | LLM can reason about code by reading it, produces fewer off-by-one and type errors | Longer lines, more tokens consumed per file | Everywhere. The token cost is worth the correctness gain |
| Explicit error handling | Prevents the most common LLM failure mode of silently swallowing exceptions | More lines of code; can feel repetitive | Any code that touches external systems, user input, or the network |
| Small files under 300 lines | LLM can see the whole context in one prompt, fewer "lost in the middle" failures | More files to manage; requires discipline to split | Every project. Break at natural boundaries |
Build it: the patterns with concrete examples
Pattern 1: Pure functions over classes
LLMs handle state badly across multiple turns. A class with mutable internal state is a trap: the LLM will forget what the state is, mutate it incorrectly, or add methods that violate invariants. Pure functions eliminate the problem entirely.
Before (class-based, stateful, high LLM error rate):
python
class PriceCalculator:
def __init__(self, tax_rate=0.08, discount_threshold=100):
self.tax_rate = tax_rate
self.discount_threshold = discount_threshold
self._applied_discounts = []
def calculate(self, items):
subtotal = sum(item.price * item.quantity for item in items)
if subtotal > self.discount_threshold:
discount = subtotal * 0.1
self._applied_discounts.append(discount)
subtotal -= discount
tax = subtotal * self.tax_rate
return subtotal + taxAn LLM maintaining this class across sessions will eventually forget to update _applied_discounts, add a method that mutates tax_rate incorrectly, or introduce a bug where calculate depends on state from a previous call.
After (pure functions, predictable, low LLM error rate):
python
def calculate_subtotal(items: list[Item]) -> float:
return sum(item.price * item.quantity for item in items)
def apply_volume_discount(subtotal: float, threshold: float = 100.0) -> float:
if subtotal > threshold:
return subtotal * 0.9
return subtotal
def calculate_tax(amount: float, tax_rate: float = 0.08) -> float:
return amount * tax_rate
def calculate_total(items: list[Item], tax_rate: float = 0.08) -> dict:
subtotal = calculate_subtotal(items)
discounted = apply_volume_discount(subtotal)
tax = calculate_tax(discounted, tax_rate)
return {
"subtotal": subtotal,
"discount": subtotal - discounted,
"tax": tax,
"total": discounted + tax,
}Every function takes explicit inputs and returns explicit outputs. The LLM cannot lose track of state because there is no state to lose. When you ask it to add a new discount rule, it adds a function to the pipeline. It does not modify a class that five other methods depend on.
Pattern 2: Descriptive names over clever ones
LLMs reason about code semantics through variable and function names. A name like x or data or tmp carries no semantic signal. The LLM has to reconstruct meaning from surrounding context, which is exactly where it makes mistakes.
Before (clever names, high error rate):
javascript
function proc(d, opts = {}) {
const res = [];
for (const x of d) {
if (x.s === "active" && x.a > (opts.min || 0)) {
const v = x.a * (opts.rate || 1.0);
res.push({ id: x.id, v });
}
}
return res.sort((a, b) => b.v - a.v).slice(0, opts.l || 10);
}Ask an LLM to modify this function, and it will guess wrong about what d, x, s, a, v, and l mean. The guesses compound: a wrong assumption about a produces a wrong modification that cascades into the sort and slice.
After (descriptive names, low error rate):
javascript
function getTopActiveAccounts(
accounts,
{ minimumBalance = 0, exchangeRate = 1.0, limit = 10 } = {}
) {
return accounts
.filter(
(account) =>
account.status === "active" && account.balance > minimumBalance
)
.map((account) => ({
id: account.id,
convertedBalance: account.balance * exchangeRate,
}))
.sort((a, b) => b.convertedBalance - a.convertedBalance)
.slice(0, limit);
}An LLM reading this function knows exactly what every variable represents. When you ask it to add a new filter for account type, it adds account.accountType === "business" instead of guessing at the field name. More tokens, yes. About 40% more. And also about 90% fewer wrong guesses on first edit.
Pattern 3: Explicit error handling
The most common LLM failure mode in production code is silent error swallowing. The LLM writes a try/except that catches Exception, logs nothing useful, and returns a default value that hides the real problem. This happens because training data is full of examples that do exactly that.
Before (implicit, LLM default, dangerous):
python
def fetch_user_data(user_id):
try:
response = requests.get(f"https://api.example.com/users/{user_id}")
return response.json()
except Exception:
return NoneThis hides every possible failure mode: network errors, auth failures, malformed responses, rate limits, server errors. Six months later, a user reports that their data is missing, and you have no idea why because the error was swallowed silently.
After (explicit, prompted, safe):
python
import requests
import logging
logger = logging.getLogger(__name__)
class UserDataError(Exception):
"""Failed to fetch user data from the API."""
def fetch_user_data(user_id: str) -> dict:
"""Fetch user data from the external API.
Raises UserDataError on any failure. Never returns None.
"""
url = f"https://api.example.com/users/{user_id}"
try:
response = requests.get(url, timeout=10)
response.raise_for_status()
return response.json()
except requests.Timeout:
logger.error("Timeout fetching user data for %s", user_id)
raise UserDataError(f"API request timed out for user {user_id}")
except requests.HTTPError as e:
logger.error(
"HTTP %s fetching user data for %s: %s",
e.response.status_code, user_id, e.response.text[:200]
)
if e.response.status_code == 404:
raise UserDataError(f"User {user_id} not found")
if e.response.status_code == 429:
raise UserDataError("Rate limited by the user API")
raise UserDataError(f"API returned HTTP {e.response.status_code}")
except requests.ConnectionError:
logger.error("Connection error fetching user data for %s", user_id)
raise UserDataError("Cannot reach the user API")To get this output from an LLM, you do not need to write the error handling yourself. You need to tell it: "Handle each error case explicitly. Log what failed and why. Raise a custom exception. Never return None on failure." The LLM knows how to write this code. It will not write it by default because its training data defaults to the silent-swallow pattern. Your prompt is what tips the balance.
Pattern 4: Small files
LLMs have finite context windows, and attention degrades toward the middle of long contexts. A 1200-line file is a trap: the LLM will make edits based on the first 200 lines and the last 200 lines, and forget what happened in the middle.
The rule: keep files under 300 lines. When a file grows past that, split it at a natural boundary. The LLM can see the whole file in one prompt without losing attention.
This is not an aesthetic preference. It is a practical constraint of the tool you are working with. A human can scroll through a 2000-line file and keep a mental model of the whole thing. An LLM cannot. The split points are the same ones you would use for a human team: one class per file, one concern per module, related utilities grouped together.
What goes wrong
| Mistake | How you notice it | The fix |
|---|---|---|
| Prompting for a class when a function would do | The LLM produces a class with mutable state, and subsequent edits introduce bugs from state inconsistency | Ask for a function. If the LLM still gives you a class, add "use pure functions, no classes, no mutable state" to the prompt |
| Using short variable names in your existing code then asking the LLM to modify it | The LLM guesses wrong about field names, types, or semantics, producing edits that look plausible but are wrong | Rename the variables before asking for edits, or provide a comment mapping short names to their meanings |
| Not telling the LLM to handle errors explicitly | The LLM defaults to except Exception: pass or returns None, hiding failures that show up weeks later | Add "handle each error case explicitly, log what failed, raise a custom exception" to every prompt that touches external systems |
| Letting files grow past 300 lines | Edits to the middle of the file are inconsistent or break things the LLM forgot about | Split the file before asking for more edits. Give the LLM the split files as context so it knows where everything lives |
| Assuming the LLM will follow your code style organically | Each new prompt produces code in a slightly different style because the LLM defaults to its training distribution | Put your style rules in CLAUDE.md. "Use type hints on all functions. Prefer list comprehensions. No single-letter variable names except loop indices." |
Confirm it worked
Take a feature you built recently and ask the LLM to rebuild it twice. First prompt: "Build a feature that does X." Second prompt: "Build a feature that does X. Use pure functions with type hints. Use descriptive variable names. Handle every error case explicitly with custom exceptions. Keep each file under 200 lines."
Compare the two outputs. The second one will have more lines, more files, and more error handling. The real test is not the initial output. Wait two days, then ask the LLM to add a new feature to each version. The version built with explicit patterns will accept the change cleanly. The version built with defaults will produce a patch that breaks something unrelated. That is the difference between code an LLM wrote and code an LLM can maintain.
Next: Prompting Patterns That Scale -- the strategies that keep a vibe-coded project coherent across days of work, not minutes.