Appearance
Model Parameters & Tuning
Every model API call comes with a handful of settings beyond the prompt text itself, and getting them wrong produces two very different symptoms: an agent that's maddeningly inconsistent from run to run, or one whose output gets cut off mid-file for no obvious reason. Neither is a prompting problem, both are parameter problems. This is short by design: there are really only three settings worth understanding well, and the rest is default values you rarely need to touch.

What you'll learn
- Temperature: why
0.0is mandatory for anything an agent needs to parse - Top-P and Top-K, and why you almost never need to tune both at once
- The truncation bug that looks like a JSON parsing bug but isn't
Temperature: how much the model is allowed to gamble
Temperature controls how the model picks its next token from the set of plausible candidates. At temperature = 0.0, it always picks the single most likely token, fully deterministic, same input, same output, every time. Turn it up toward 1.0 and the model starts picking from a wider spread of plausible-but-not-most-likely tokens, which is what gives creative writing its variety and unpredictability.
For an agent that needs to produce code, JSON, or anything else a program downstream will parse, the choice is easy: temperature = 0.0. Anything above that risks a syntactically valid-looking response that's actually got a stray comma or a mismatched bracket, because the model gambled on a slightly-less-likely token at the wrong moment. Save higher temperatures for tasks where variety is the point, brainstorming, roleplay, creative copy, never for anything a parser has to read afterward.
Top-P and Top-K: narrowing the field before temperature even applies
These two settings filter the candidate pool before temperature scaling happens, and you'll rarely need to tune both at once:
- Top-K caps the candidate pool to a fixed number ,
top_k = 40means only the 40 most likely next tokens are even considered, no matter how the probabilities are distributed. - Top-P (nucleus sampling) is more adaptive: it keeps adding candidates, most-likely first, until their combined probability crosses a threshold ,
top_p = 0.95keeps whatever set of tokens covers 95% of the likely probability mass, which might be 5 tokens or 50 depending on how confident the model is.
In practice: if you've already set temperature = 0.0, top-p and top-k barely matter, since the model is only picking the single top token anyway, leave them at their defaults. If you're deliberately running at a higher temperature for a reasoning or brainstorming task, top_p = 0.95 is a reasonable way to keep some variety while still pruning obviously bad low-probability tokens.
Max output tokens: the limit that silently breaks JSON parsing
This one causes more confusing bugs than it should. max_output_tokens caps how much the model is allowed to generate in a single response, and if you set it too low for the task (say, 256 tokens for something that needs to generate a few hundred lines of code), the response gets cut off mid-token, mid-line, mid-anything. The result looks exactly like a parsing bug: a JSONDecodeError, invalid Python syntax, a truncated file. It isn't a parsing bug. It's a limit set too tight.
What goes wrong
| Mistake | How you notice it | The fix |
|---|---|---|
| Structured output occasionally fails to parse | Intermittent JSONDecodeError, works most of the time | Set temperature = 0.0 for anything parsed downstream |
| Output cuts off mid-file or mid-sentence | Truncated code, unterminated JSON, missing closing brackets | Raise max_output_tokens , check the actual length of a full expected response first |
| Creative/brainstorming output feels flat and repetitive | Every run produces nearly identical phrasing | Raise temperature for tasks where variety is actually wanted, don't leave everything at 0.0 |
| Tuning top-p and top-k together with no clear reason | Unpredictable behavior that's hard to reason about | Pick one (usually top-p) and leave the other at its default, tuning both at once rarely adds value |
Confirm it worked
To verify that your temperature setting produces stable, deterministic results, run this script to invoke the model multiple times and assert that the response format remains completely consistent:
python
# Run the same structured-output prompt 20 times at temperature 0.0
# and confirm every single response parses cleanly, if even one fails,
# the problem is the prompt's structure, not the temperature setting.
results = [call_model(prompt, temperature=0.0) for _ in range(20)]
assert all(is_valid_json(r) for r in results)Next: Day 2 , Agent Tools & Interoperability, where your agent starts reaching real APIs and executing real code.