Beyond Systems

LLM API Guide

Build production-ready AI applications: API fundamentals, prompt engineering, rate limiting, cost optimization, and scalable patterns for LLM-powered services.

Category: AI/API Status: 4 Topics

Lessons

Master the practical skills to build and deploy LLM-powered applications at scale.

LLM API Fundamentals BeginnerEssential +
ProviderModelPricingKey Features
OpenAIGPT-4, GPT-3.5Pay-per-tokenMost mature, extensive docs, function calling
AnthropicClaudePay-per-tokenSafer, longer context windows
GoogleGeminiPay-per-tokenMultimodal, function calling
CohereCommandPay-per-tokenEmbeddings, reranking
DeepSeekDeepSeek-ChatLow-costOpen-source friendly

Basic API Call Structure

Most LLM APIs follow a similar pattern. Here's a complete example with OpenAI:

import openai, os

client = openai.OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain quantum computing in simple terms."}
    ],
    temperature=0.7,
    max_tokens=500
)
print(response.choices[0].message.content)

Temperature and Sampling

ParameterRangeEffect
temperature0.0 - 2.0Creativity vs. determinism (lower = more predictable)
top_p0.0 - 1.0Nucleus sampling for diversity
max_tokens1 - 4096Response length limit
frequency_penalty0.0 - 2.0Reduce repetitive responses

Error Handling

from openai import OpenAI, OpenAIError
import time

def call_llm_safely(client, messages, max_retries=3):
    for attempt in range(max_retries):
        try:
            response = client.chat.completions.create(
                model="gpt-4o", messages=messages, temperature=0.7
            )
            return response.choices[0].message.content
        except OpenAIError as e:
            if attempt < max_retries - 1:
                time.sleep(2 ** attempt)
            else:
                raise e
System Prompts: Setting the Stage for AI Behavior Beginner +

System prompts are the invisible director setting expectations. They run before all conversation, shaping how the model responds. Without them, models default to generic helpfulness. With them, you get consistent, reliable outputs.

Why they matter: imagine asking a junior developer to summarize quarterly earnings. Without context, they'd ask for spreadsheet access. With "summarize email reports from finance@company.com sent before March 31," they ship the right work in minutes. System prompts replace dozens of follow-up questions.

Anatomy of a Good System Prompt

You are a senior ML engineer specializing in inference optimization.
Provide concise, production-focused answers.
Always explain tradeoffs.
Never mention training procedures unless asked.
Format code in triple backticks.
Respond to technical questions in under 200 words.

This prompt establishes capabilities, constraints, and tone. The model knows its role, limits, and audience.

System Prompts vs User Prompts

A user prompt with "summarize this" gets a summary. A user prompt "summarize using these four headings" still works, but a system prompt "you are a technical writer" makes summaries consistently structured. System prompts set the stage, user prompts provide the script.

Common Patterns

  • Role-based: "You are a Python expert..." → direct, practical answers with code examples.
  • Constraint-based: "Respond in under 100 words..." → concise output without follow-up length concerns.
  • Format-based: "Always use markdown..." → consistent formatting for automated parsing.
  • Personality: "Be direct but friendly..." → tone without sacrificing clarity.

Real-World Example: API Client Helper

You are an API integration specialist.
Write production-ready Python using httpx and type hints.
Include proper error handling and docstrings.
Never include API keys. Use placeholders.

User prompt: "Create a client for OpenWeatherMap" → result: code with retry logic, proper type definitions, and clear error messages, without you needing to specify those requirements.

Dynamic System Prompts

Current task: Debug slow database queries
Context: PostgreSQL, 1M+ rows, queries over 5 seconds
Focus: Query optimization, indexing, EXPLAIN ANALYZE

Now the model shifts from general advice to specific debugging strategies.

Testing System Prompts

Bad system prompts produce inconsistent results. Test by asking the same question 3 times. If answers vary wildly, the prompt lacks clarity or constraints. Good system prompts yield predictable responses across inputs, consistency enables automation.

Practical Exercise

Create a system prompt for code review:

You are a senior Python developer reviewing pull requests.
Focus on security, performance, and maintainability.
Point out specific lines with line numbers.
Suggest concrete alternatives.
Remain constructive and professional.

Test with this diff:

- if user_input:
+ if len(user_input) > 0:

A good system prompt should flag the unnecessary len() call as a performance anti-pattern.

Key takeaways

System prompts are your first line of quality control. They define who the model is, what constraints apply, and how responses should look. Spend time crafting them, they're worth more than perfect user prompts.

Structured Outputs: Turning Freeform Text Into Predictable Data Intermediate +

LLMs excel at unstructured text but stumble when you need reliable data. Ask for "a list of features" and get prose. Ask for JSON and sometimes still get "Sure, here's the JSON: {...}" wrapped in commentary. Structured outputs turn probabilistic text back into deterministic code.

The Problem: Unreliable Output

User prompt: "Give me three Python features" produces a paragraph of prose. Three features? Hard to parse. Should you extract with regex? What happens when the format changes next time?

The Solution: Constrained Output

[
  {"name": "list comprehensions", "description": "..."},
  {"name": "decorators", "description": "..."},
  {"name": "type hints", "description": "..."}
]

The LLM produces valid JSON. Your code parses it. No guesswork.

How Structured Outputs Work

response_format = {"type": "json_object"}

# or, for more control:
response_format = {
    "type": "json_schema",
    "json_schema": {
        "name": "feature_list",
        "schema": {
            "type": "object",
            "properties": {
                "features": {
                    "type": "array",
                    "items": {
                        "type": "object",
                        "properties": {
                            "name": {"type": "string"},
                            "description": {"type": "string"}
                        },
                        "required": ["name", "description"]
                    }
                }
            }
        }
    }
}

Real-World Example: API Specification

Without structured output, "design a REST API for a todo app" gets you prose you copy-paste and debug. With it:

schema = {
    "type": "object",
    "properties": {
        "endpoints": {
            "type": "array",
            "items": {
                "type": "object",
                "properties": {
                    "method": {"type": "string"},
                    "path": {"type": "string"},
                    "description": {"type": "string"},
                    "body": {"type": "object"}
                }
            }
        }
    }
}

Output imports directly into your FastAPI routes. Zero manual parsing.

Common Patterns

Batch extraction: pull structured data from unstructured text (e.g. support emails → {"name", "email", "issue"}).

Template validation: ensure the model fills templates correctly.

Configuration generation: turn natural language into config objects, "make this AWS Lambda config" → structured JSON parameters.

Trade-offs

Pros: deterministic output, no parsing logic, type-safe downstream processing.

Cons: rigid format, less natural language flexibility, schema definition effort.

Use structured outputs when you need reliability. Use freeform when you need explanation.

Getting Started

import openai, json

response = openai.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "List 3 Python features as JSON array: {name, description}"}],
    response_format={"type": "json_object"}
)
features = json.loads(response.choices[0].message.content)

Key takeaways

Structured outputs are guardrails, not cages. They force LLMs into the format you need, eliminating brittle parsing logic. Define your schema once, get perfect data every time.

Function Calling: Making LLMs Call Your Code Advanced +

LLMs are great at conversation but terrible at running code. Function calling bridges this gap, it lets the model decide what to do, while you write how to do it.

Without function calling: you write "call search function with this query" in your prompt, the model hallucinates search parameters, your code doesn't run. With function calling: the model returns structured function arguments and your code executes them directly.

How It Works

You provide two things to the API: function definitions (the tools available) and messages (the conversation). The model replies with either text or a function call object. When it calls a function, you execute it and send the result back, the model continues reasoning.

Simple Example: Weather Lookup

functions = [{
    "name": "get_weather",
    "description": "Get current weather for a city",
    "parameters": {
        "type": "object",
        "properties": {
            "city": {"type": "string", "description": "City name, e.g. San Francisco"},
            "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "default": "celsius"}
        },
        "required": ["city"]
    }
}]

# User asks: "What's the weather in Tokyo?"
# Model responds:
# {"name": "get_weather", "arguments": "{\"city\": \"Tokyo\", \"unit\": \"celsius\"}"}

def get_weather(city: str, unit: str = "celsius") -> str:
    return f"{city}: 22°{unit[0].upper()}"

Complex Example: Restaurant Booking

functions = [{
    "name": "book_table",
    "description": "Book a restaurant reservation",
    "parameters": {
        "type": "object",
        "properties": {
            "restaurant": {"type": "string"},
            "party_size": {"type": "integer", "minimum": 1, "maximum": 20},
            "datetime": {"type": "string", "format": "date-time"},
            "occasion": {"type": "string", "enum": ["birthday", "anniversary", "date", "none"]}
        },
        "required": ["restaurant", "party_size", "datetime"]
    }
}]

# "Book a table for 4 at Sushi House tomorrow at 7 PM for my anniversary" ->
# {"name": "book_table", "arguments": {"restaurant": "Sushi House", "party_size": 4,
#   "datetime": "2024-01-15T19:00:00", "occasion": "anniversary"}}

When to Use Function Calling

Use it when: the model needs to make decisions, external data is required, you need reliable input extraction, or it's part of an automation workflow.

Don't use it when: it's pure conversation, a simple calculation the model can just reason through, or the information is already in context.

Building a Complete Loop

import openai, json

available_functions = {
    "get_weather": lambda city, unit="celsius": f"{city}: 22°",
    "book_table": lambda restaurant, party_size, datetime, occasion="none": "Booked!"
}

def run_conversation(messages):
    while True:
        response = openai.chat.completions.create(
            model="gpt-4o", messages=messages, functions=[...], function_call="auto"
        )
        message = response.choices[0].message
        if message.get("function_call"):
            func_name = message["function_call"]["name"]
            func_args = json.loads(message["function_call"]["arguments"])
            result = available_functions[func_name](**func_args)
            messages.append({"role": "assistant", "content": result, "function_call": message["function_call"]})
        else:
            return message["content"]

messages = [{"role": "user", "content": "What's the weather in London?"}]
print(run_conversation(messages))

Common Patterns

  • Orchestrator pattern: LLM decides which functions to call and in what order, great for multi-step workflows.
  • Extract pattern: LLM extracts structured data from user input into function arguments, clean separation of NLU and execution.
  • Validation pattern: model calls validation functions to check if proposed solutions work.

Tool Use vs Function Calling

Function calling is simpler but less flexible. Tool use (used by agents) supports multiple parallel calls, tool descriptions, and more complex interactions. Start with function calling, move to tool use when you need multi-step reasoning.

Getting Started Checklist

  • Define function signatures for your APIs
  • Map model calls to actual implementations
  • Handle errors (invalid arguments, missing functions)
  • Return results in conversational format
  • Test with ambiguous and complex queries

Key takeaways

Function calling turns LLMs into decision engines while keeping your code execution reliable. Define clear interfaces, handle the loop, and let the model orchestrate your tools.