Skip to main content
← Back to Blog
2026-10-03 · 9 min readDev ToolsAI Automation

ChatGPT for Developers: A Practical Guide to Building Better Workflows

From the ChatGPT app to the OpenAI Responses API: prompt patterns, structured JSON output, a small Python pipeline you can reuse, and the guardrails I keep around generated code.

Avnish Yadav
Avnish Yadav
Developer & Automation Builder
ChatGPT for Developers: A Practical Guide to Building Better Workflows

For a long time I used ChatGPT the way most developers start: paste an error, get a fix, move on. It works, but it does not compound. Every session starts from zero, the quality depends on how tired I am when I write the prompt, and nothing I learn ends up in code. The shift that made it useful was treating it as two separate tools: the ChatGPT app for thinking and exploring, and the OpenAI API for anything I want to run more than twice.

This guide covers both sides: prompt patterns that hold up for code work, the current Responses API in Python, structured JSON output, a small pipeline you can reuse, and the places where I do not use it at all. Model names and API details are from OpenAI's documentation as of October 3, 2026 (linked in the sources); check them again before you copy anything into production, because they change often.

ChatGPT app or OpenAI API?

Both run OpenAI models, but they solve different problems.

Use the ChatGPT app for exploration: understanding an unfamiliar library, sketching an approach, rubber-ducking a bug, drafting a README you will edit anyway. The conversation is the product, and you are in the loop for every step.

Use the API when the same task will run again: a code review step in CI, summarising issues every morning, generating test fixtures, triaging logs. The API gives you control over instructions, input, output format and model version, and the result lands in your code instead of a chat window.

My rule of thumb: if I catch myself pasting the same kind of prompt into the app for the third time, it becomes a script.

Prompt patterns that hold up for code

These work the same in the app and through the API.

Give the model a role and a contract

A short set of standing instructions removes most vague answers. For code review I use something like this:

You are a senior Python reviewer. Review the code below for bugs, security issues,
performance problems and readability, in that order of priority.
For each issue give: severity (high, medium, low), the line or function, the problem,
and a concrete fix. If you are not sure something is a bug, say so. Do not rewrite
the whole file.

The last two sentences matter most. Asking the model to flag uncertainty and to stay narrow cuts down on confident, invented problems and on huge rewrites you then have to diff by hand.

Describe the data, not just the task

"Write a function to parse JSON" gets you a generic function. Describing the actual shape gets you one you can use:

Write a Python function that takes the JSON body of GET /users from our API.
The body is {"data": [{"name": str, "email": str, "role": "admin" | "member"}], "next": str | null}.
Return a list of dicts with name, email and role. Raise ValueError if "data" is missing.
Use only the standard library and add type hints.

Include the input shape, the output shape, the error behaviour and the constraints (libraries, Python version, style). The more of the contract you state, the less you rewrite.

Ask for tests next to the code

When I ask for a function I also ask for three or four pytest cases, including one edge case. Reading the tests is the fastest way to see whether the model understood the problem, and running them is the fastest way to find out whether it did not.

Keep conversations short and scoped

Long chats drift. Earlier assumptions linger, and the model starts mixing up versions of the same file. When a thread changes topic, I start a new one and paste only what is relevant. Through the API this is natural: each call gets exactly the context you send.

For more on prompt structure, see Crafting effective prompts for LLM performance.

Calling the OpenAI API from Python

OpenAI's quickstart uses the Responses API. Install the SDK and set your key; the client reads OPENAI_API_KEY from the environment.

pip install openai
export OPENAI_API_KEY="<YOUR_OPENAI_API_KEY>"
from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-6.1-sol",
    instructions="You are a concise senior Python reviewer. Flag uncertainty explicitly.",
    input="Review this function:\n\ndef avg(xs):\n    return sum(xs) / len(xs)",
)
print(response.output_text)

Which model ID?

OpenAI's models page currently lists three flagship models:

Model ID OpenAI's description
GPT-6 Astra gpt-6-astra Most capable, for the most demanding work
GPT-6.1 Sol gpt-6.1-sol Near-Astra performance for complex work at a lower cost
GPT-6 Luna gpt-6-luna Most efficient, for focused, high-volume tasks

OpenAI recommends GPT-6 Astra if you are not sure where to start. For day-to-day developer automation I start with Sol and drop to Luna for simple, high-volume steps such as classification or short summaries. The pricing page lists them at $10 / $50 (Astra), $2 / $10 (Sol) and $0.10 / $0.50 (Luna) per million input / output tokens for standard short-context requests, so the choice is worth testing rather than defaulting to the biggest model.

A few API details that trip people up:

  • instructions applies to one request only. OpenAI's docs note that if you chain turns with previous_response_id, the instructions from earlier turns are not carried over. Send them on every call.
  • Reasoning effort replaces temperature tuning. The GPT-6 models are reasoning models, controlled with reasoning={"effort": ...}. Supported values depend on the model (for example, gpt-6.1-sol does not accept "none" or "minimal"). Older advice to "set temperature to 0.1 for code" does not map cleanly onto reasoning models, so check the model's page before relying on sampling parameters.
  • Pin a snapshot in production. OpenAI recommends pinning production apps to a dated model snapshot so behaviour does not shift under you. Use the alias while experimenting and the snapshot once a workflow is live.

Structured outputs: JSON you can trust

Most automation breaks at the boundary between the model's text and your code. Structured Outputs fixes that: you give a schema and the response conforms to it. With Pydantic, the SDK's parse helper does the work:

from openai import OpenAI
from pydantic import BaseModel

client = OpenAI()

class ReviewIssue(BaseModel):
    severity: str          # "high" | "medium" | "low"
    location: str
    problem: str
    fix: str

class Review(BaseModel):
    issues: list[ReviewIssue]
    summary: str

response = client.responses.parse(
    model="gpt-6.1-sol",
    instructions="You are a senior Python reviewer. Flag uncertainty explicitly.",
    input=open("app/billing.py").read(),
    text_format=Review,
)
review = response.output_parsed
for issue in review.issues:
    print(f"[{issue.severity}] {issue.location}: {issue.problem}")

If you do not use Pydantic, pass a JSON Schema in text={"format": {"type": "json_schema", "name": "...", "schema": {...}, "strict": True}}. Either way, the schema guarantees shape, not truth. A well-formed issue can still be wrong, so the output is a suggestion a human or a test confirms.

A repeatable pipeline you can adapt

Here is the pattern I reuse for most small automations: input, model call with a fixed contract, structured output, storage. This example summarises text (an issue, a meeting note, an article) into a SQLite table.

import sqlite3
from openai import OpenAI
from pydantic import BaseModel

client = OpenAI()
MODEL = "gpt-6-luna"  # cheap, high-volume step; move up only if quality is not enough

class Digest(BaseModel):
    summary: str
    key_points: list[str]
    tags: list[str]

def digest(text: str) -> Digest:
    response = client.responses.parse(
        model=MODEL,
        instructions=(
            "Summarise the input for a developer in 2-3 sentences, list up to 5 key "
            "points, and give 1-4 lowercase tags. Use only facts in the input."
        ),
        input=text,
        text_format=Digest,
    )
    return response.output_parsed

def store(db: sqlite3.Connection, source: str, d: Digest) -> None:
    db.execute(
        "INSERT INTO digests (source, summary, key_points, tags) VALUES (?, ?, ?, ?)",
        (source, d.summary, "\n".join(d.key_points), ",".join(d.tags)),
    )
    db.commit()

Swap the Digest model and the instructions and the same skeleton triages GitHub issues, drafts release notes from commit messages, or generates test case ideas from a spec. When a step needs tools and decisions rather than one call, it is time for an agent loop; I cover that in Building AI agents from scratch. If you would rather wire this up visually, n8n can call the same API; my n8n automation guide is a good starting point.

Log triage and documentation

Two uses that have stayed in my own workflow:

Log triage. A small script takes the tail of an error log, groups repeated stack traces, and asks the model for a likely cause and a suggested fix per group, as structured output. It writes the result to a Markdown file I read in the morning. I deliberately stopped short of letting it open pull requests on its own. The suggestions are a starting point for my debugging, not a replacement for it, and I would rather review a list than review an unexpected PR.

Documentation. After the code is written and tested, I ask for docstrings and a README section from the actual source, with a fixed template (purpose, usage example, parameters, errors). Because the code exists first, the docs describe what it does, not what I meant it to do. I still read every line, because a confident docstring for the wrong behaviour is worse than none.

Mistakes I made early on

  • Treating output as correct by default. Generated code can look right and be wrong in subtle ways: an off-by-one, a missing await, a deprecated API. Run it, test it, read it.
  • Sending too much context. Pasting a whole repository into a chat makes answers vaguer, not better. Send the function, its callers and the error.
  • Not stating constraints. If you need standard library only, Python 3.11, or no new dependencies, say so up front.
  • Letting the model choose the output format. For anything a program reads, use Structured Outputs instead of asking for JSON in prose.
  • Hard-coding a model name everywhere. Keep it in one config value so trying a new model is a one-line change.

When not to use ChatGPT

  • Hard real-time paths. A model call adds network and generation time you do not control. Keep it out of request paths with tight latency budgets; run it asynchronously instead.
  • Exact, deterministic logic. Calculations, permissions, billing rules and anything that must give the same answer every time belong in normal code. The model can help you write that code; it should not be that code.
  • Sensitive data without a policy. OpenAI's API docs state that data sent to the API is not used to train its models unless you opt in, and that abuse-monitoring logs are kept for up to 30 days by default. That is the API; the ChatGPT app has its own settings that depend on your plan and workspace. Know which one you are using and what your client contracts allow before you paste production data or secrets.
  • Unpinned critical workflows. Models are updated. If a business process depends on a specific behaviour, pin a snapshot and keep a test set to rerun before you upgrade.

Frequently asked questions

Can ChatGPT replace a developer?

No. It speeds up boilerplate, first drafts, explanations and debugging hypotheses, but architecture, trade-offs, security and knowing what to build stay with you. The developers who get the most from it are the ones who can tell when its answer is wrong.

Which OpenAI model should I use for coding tasks?

OpenAI suggests GPT-6 Astra (gpt-6-astra) as the starting point for complex reasoning and coding. For automated, repeated steps I test GPT-6.1 Sol and GPT-6 Luna first because they cost less, and only move up when my test cases need it.

Is it safe to use the API with private code?

OpenAI states that API data is not used for training unless you opt in, with abuse-monitoring logs kept for up to 30 days by default. Whether that is acceptable depends on your employer's or client's rules. Never send secrets or credentials, and strip them from logs before sending.

How do I make the output consistent?

Use clear standing instructions, describe the input and output precisely, add one or two examples, and use Structured Outputs for anything your code reads. Pin a model snapshot for production workflows.

Can I automate ChatGPT without writing code?

Yes. Tools like n8n can call the OpenAI API from a workflow. For anything that needs validation, tests or version control, I still prefer a small script like the pipeline above.

Sources

Verified against the sources below on October 3, 2026. Products and docs change often: check the linked sources if something looks different.

  1. OpenAI API: Models
  2. OpenAI API: Quickstart
  3. OpenAI API: Text generation (instructions, snapshots)
  4. OpenAI API: Structured Outputs
  5. OpenAI API: Reasoning models
  6. OpenAI API: Pricing
  7. OpenAI API: Data controls
Share
Discussion

Comments

Loading comments...

Add a comment

Comments are reviewed before they appear.