Use case

OpenAI Dots for Developers: Feedback to Tested Pull Requests

OpenAI's developer example is a Dot that watches recurring feedback, scopes small improvements and bug fixes, builds and tests them, and brings back complete pull requests with videos.

Last verified: September 30, 2026 · Reviewed by AstraDot editorial desk

Engineering backlogs fill up with small, well-understood work: the bug that keeps being reported, the rough edge customers mention in every call, the fix everyone agrees is right but nobody has time to write. It is rarely hard; it is just relentless. OpenAI's developer scenario for Dots targets exactly that category of work, and it is worth understanding precisely because it is modest rather than dramatic.

This guide walks through the scenario as OpenAI described it, then separates the parts a Dot can carry from the parts that should stay with your team. If you want the product basics first, start with what Dots are.

The scenario

A Dot watches recurring customer feedback. Instead of treating each report as an isolated ticket, it looks for the pattern: the same complaint arriving from different accounts, the small improvement that would quiet several threads at once. It scopes those improvements and bug fixes, builds and tests them, and brings back complete pull requests — with videos that show the work.

The unit of value is not "code written." It is the full loop: noticing a signal, deciding it is small enough to act on, implementing it, testing it, and handing it over in a form a reviewer can accept or reject quickly.

What the Dot does end-to-end

  • Watches feedback. It tracks recurring customer feedback over time and learns from it, so recurring issues rise above noise.
  • Scopes the change. It turns a pattern into a specific, bounded improvement or bug fix rather than a vague backlog item.
  • Builds and tests. With its own cloud computer and browser, it can write and test code directly.
  • Packages the pull request. It returns a complete PR with videos, so the reviewer sees the change and its behaviour, not just a diff.
  • Keeps going. A Dot works 24/7 and can handle several projects at once, so this loop continues alongside other work.

Context carries across channels. You can reach the Dot through ChatGPT on desktop, web or mobile, or in Slack and Teams, and it can message you when something needs attention. Texting via iMessage/RCS is coming on a limited Pro waitlist.

What stays with you

The scenario is deliberately bounded, and the boundary is where engineering judgement lives.

  • Review and merge. The Dot brings pull requests; people review and merge them.
  • Architecture. Deciding how a system should be shaped is not the Dot's call.
  • Security and data handling. Choices about credentials, access and sensitive code remain human decisions.
  • Scope judgement. What counts as a small fix worth doing, and what should be a design conversation, is yours to set.

The controls reinforce this. Auto-review checks the Dot's actions against your instructions, Custom Rules and safety requirements before consequential steps. Proactive research uses read-only tools and cannot send messages, change app content, or control your browser or computer. Some sensitive tasks always stay with you. You can follow progress in the activity view, and OpenAI can pause or stop a Dot if monitoring finds a concern. Full details are in our security and privacy guide.

How to split the work

A useful division of labour is to give the Dot the repetitive middle of a well-defined loop and keep the decisions at both ends.

StepDotYour team
Surface feedbackWatch and cluster recurring reportsDecide which themes matter
Scope the fixDraft a bounded changeApprove scope and approach
ImplementWrite the codeSet conventions and architecture
TestRun tests and record a videoDefine what "passing" means
ShipOpen a complete PRReview, merge, release

Risks and guardrails

The main risks are familiar from any automation: the Dot can misread intent, over-generalise from a few complaints, or produce a change that passes tests but misses context. Dots can make mistakes, so consequential work needs review.

  • Keep the blast radius small. Start with low-risk repositories and reversible changes.
  • Write Custom Rules. Encode coding standards, forbidden areas and required reviewers so auto-review has something concrete to check against.
  • Require the video. Treat the recorded evidence as part of the deliverable, not a nicety.
  • Watch the activity view. Background work should be visible, and surprises are a signal to narrow scope.
  • Track usage. Conversations with a Dot do not count toward ChatGPT usage limits, but tasks started in Codex or ChatGPT Work do.

A practical rollout checklist

  1. Pick one recurring, low-risk class of feedback with a clear owner.
  2. Connect the systems the Dot needs and confirm what it can reach.
  3. Name the workflow in your primary Dot and describe the outcome you want.
  4. Write Custom Rules for standards, restricted areas and reviewer requirements.
  5. Agree that every pull request gets a human review and merge.
  6. Measure two things for a month: review time per PR and how many PRs you accept.
  7. Widen scope only after the loop is boringly reliable. The Agent Readiness Scorecard helps you judge the next candidate.

Source: OpenAI, Introducing dots and the Dots help center.

Frequently asked questions

Does a Dot merge its own pull requests?
No. The scenario OpenAI described has the Dot bring back complete pull requests with videos, not merge them. Review and merge stay with your team, alongside architecture and security decisions. See Dots security and privacy.
How does the Dot know what feedback matters?
It watches recurring customer feedback and learns from feedback over time, so patterns surface instead of one-off tickets. You still define what counts as in scope through your instructions and Custom Rules.
Can it run tests and build code?
Yes. Each Dot has its own cloud computer and browser and can write and test code, which is what makes an end-to-end fix-and-test loop possible. Tasks started in Codex or ChatGPT Work do count toward ChatGPT usage limits.
Where should we start?
Pick one narrow, recurring problem with a clear reviewer and a fast test signal. Our free Agent Readiness Scorecard can help you judge whether a workflow is ready for an always-on agent.