Use case

OpenAI Dots for Scientists: Analysis That Updates With New Evidence

OpenAI's scientist example: a Dot that learns your research question and how you evaluate evidence, reruns analyses as data arrives, investigates surprises, and flags what needs review.

Last verified: September 30, 2026 · Reviewed by AstraDot editorial desk

Research is iterative by nature. Data arrives over days or weeks, an analysis that was correct last month is out of date today, a figure has to be regenerated, and an unexpected result demands a closer look before it means anything. Much of that loop is mechanical reruns and bookkeeping wrapped around a small number of judgements. OpenAI's scientist scenario for Dots occupies exactly that space.

For the product basics, start with what Dots are. This page covers the research workflow, the reproducibility and review practices that make it trustworthy, and the parts that should never be automated.

The scenario

A scientist's Dot learns the research question and how the researcher evaluates evidence. It reruns analyses as data arrives, investigates unexpected results, updates figures, and revises explanations — while flagging what needs review rather than settling it. The emphasis on "flagging" is deliberate: the Dot keeps the analytical machinery current so the researcher can concentrate on interpretation.

What the Dot does end-to-end

  • Learns the question. It picks up both what you are asking and the standards by which you accept evidence.
  • Reruns analyses. When new data arrives, it repeats the analysis rather than leaving a stale result in place.
  • Investigates surprises. Unexpected results get a closer look instead of being silently smoothed over.
  • Updates figures. Outputs are regenerated so the visuals match the current data.
  • Revises explanations. Written interpretations are adjusted to reflect what the analysis now shows.
  • Flags for review. Anything that needs a human judgement is surfaced rather than decided.

Every Dot is powered by GPT-6 Astra, has its own cloud computer and browser, can write and test code, works 24/7, and can handle several projects at once. For research that matters because reruns can happen on their own schedule, not only when someone remembers. Reach the Dot through ChatGPT, Slack and Teams, and it can message you when something needs attention; texting via iMessage/RCS is coming on a limited Pro waitlist.

Reproducibility and review

Rerunning an analysis is only useful if you can trust and repeat it. A few practices make the difference:

  • Version the inputs. Point the Dot at datasets that are themselves versioned so a rerun is traceable.
  • Keep the scripts. Because the Dot can write and test code, favour stored, inspectable scripts over ad-hoc steps.
  • Make the review boundary explicit. The Dot flags; a researcher confirms. Auto-review checks the Dot's actions against your instructions, Custom Rules and safety requirements before consequential steps.
  • Watch the activity view. Background work should be visible, and the record of what changed is part of the evidence.

Proactive research uses read-only tools, so a Dot cannot change app content or control your browser or computer on its own initiative. Some sensitive tasks always stay with you, and OpenAI can pause or stop a Dot if monitoring finds a concern. The security and privacy guide covers the controls; the model's behaviour is documented in the GPT-6 Astra system card.

What should never be automated

The scenario keeps a clear line between calculation and judgement. These stay human:

ActivityWhy it stays with the researcher
Interpreting findingsMeaning depends on context the Dot cannot weigh
Judging validityMethodological and domain expertise is required
Deciding what to publishPublication is an accountable human act
Ethical and consent decisionsThese are governance calls, not computations
Final figures and claimsNothing goes out without a person checking it

Put simply: a Dot can keep the analysis current and flag what changed, but it should not conclude what the evidence means. Dots can make mistakes, and in research the cost of an unexamined error is high, so consequential output needs review.

Risks and guardrails

  • Silent drift. An updated figure can move without anyone noticing; require a visible change log.
  • Over-fitting to a familiar answer. State the evaluation standard up front so the Dot does not chase a preferred result.
  • Data boundaries. Keep sensitive or restricted data within approved systems and review what the Dot can reach.
  • Automation bias. Because reruns are convenient, it becomes tempting to accept them; keep the review step real.

A practical rollout checklist

  1. Choose one repetitive analysis with a clear, testable definition of done.
  2. Version the datasets and scripts the Dot will use.
  3. Describe the research question and your evidence standard in the Dot's instructions.
  4. Write Custom Rules that reserve interpretation and publication for people.
  5. Require a visible change log whenever a figure or explanation updates.
  6. Review early reruns closely before trusting the loop.
  7. Expand only once reproducibility holds. The Agent Readiness Scorecard helps you pick the next step.

Source: OpenAI, Introducing dots and the GPT-6 Astra system card.

Frequently asked questions

Should a Dot decide what a result means?
No. Scientific judgement should stay with the researcher. In OpenAI's scenario the Dot reruns analyses, investigates unexpected results, updates figures and revises explanations, but it flags what needs review rather than concluding. Read Dots security and privacy.
Can the analysis be reproduced later?
That is a matter of how you set up the Dot and your tooling. Because each Dot has its own cloud computer and can write and test code, favour versioned inputs and scripts so a rerun can be repeated and inspected.
What model runs a Dot used for research?
Every Dot is powered by GPT-6 Astra. The model page covers GPT-6 Astra, and the system card is the official source for safety and evaluation detail.
Where is the safest place to start?
Begin with a repetitive analysis step that already has a clear, testable definition of done, such as regenerating a standard figure when a dataset updates. Our free Agent Readiness Scorecard can help you judge readiness.