
Recently, while using Dot, we received this message:
Your dot is on a break. Congrats you are in the very top users and have hit our abuse prevention limit. Check back in a bit!
Dot was working on an outbound email campaign that was already running on a daily cron, a recurring schedule. We were on Max, and it took several hours before we could resume. Those are the details from our experience; we did not capture a usage counter or an exact reset time.
For a business owner, the problem is practical. You have work underway and a schedule to keep. “Check back in a bit” gives you very little to plan around.
That experience raises two questions: what might Dot be limiting, and how would the same work behave if we controlled the inference capacity?
What we know about the Dot limit
The message identifies an abuse-prevention limit. It does not state how usage is counted, the threshold, the reset schedule, or which activity triggered it. It also does not establish that our campaign violated a policy or that every background task stopped.
OpenAI's Dot documentation distinguishes conversations with a Dot from tasks delegated to Work or Codex, which use those products' allowances. It also describes an allowance for deeper work. The documentation we checked on October 8, 2026 did not give a numerical threshold or cooldown for the specific message we received.
Our “Max” plan label does not identify the allowance involved. We would need account-specific information to connect this event to a published plan limit.
Our best guess: a usage window or an activity guardrail
Our leading hypothesis is that sustained agent work reached a usage guardrail measured over a window of time. The several-hour interruption is compatible with that explanation. It does not tell us whether the window rolls continuously, resets at a fixed time, or uses some other release condition.
The existing daily schedule adds useful context. One scheduled trigger can start many operations, and successive runs could contribute to a longer usage window. A larger batch or more retries could also make one day's work more demanding. These are possible mechanisms, not things we observed in the logs. A recurring schedule does not, by itself, establish a fixed amount of processing per run.
An agent can perform substantial work between messages. Researching companies, reading context, drafting email, checking results, and retrying steps could each consume resources. A small number of instructions from the owner can therefore represent a much larger amount of processing. We do not know which of these actions, if any, counted toward this limit.
The outbound campaign makes an activity-specific control another plausible explanation. Here is how we would distinguish the possibilities:
| Hypothesis | Why it could fit | What remains unknown |
|---|---|---|
| A usage allowance over time | Sustained work followed by a multi-hour interruption is compatible with a cooldown | The unit of usage, threshold, and reset rule |
| A burst or concurrency limit | An agent may create many requests or run several tasks together | Whether our campaign actually produced that pattern |
| An outbound-activity guardrail | The task involved email outreach, and the message names abuse prevention | Whether the control assessed sending activity, general agent activity, or something else |
These are hypotheses, not findings about OpenAI's implementation. There is no defensible conversion from this screenshot to “emails per day,” “messages per hour,” or “tokens per week.”
To narrow it down, we would record when the scheduled run started, its batch size and retries, the last successful action, the first blocked action, other active tasks, any displayed usage information, and the first successful attempt after recovery. Those observations could support a useful support request. A second event would still need interpretation; it would not automatically reveal a universal quota.
The business cost is uncertainty about when work can continue
A rate limit can be workable when the team knows its size and reset time. You can schedule a batch, assign urgent work first, or decide that a job needs more capacity.
An unexplained interruption makes those decisions harder. Should the team wait, take over manually, or move the task elsewhere? Can tomorrow's campaign finish before the sales team starts its day? With a daily schedule, the team also needs to know whether interrupted work resumes, is skipped, or is attempted again. Our screenshot does not establish what happened to the scheduled job itself.
The subscription price captures only part of that cost. Staff time spent checking progress, reconstructing work, or handling missed handoffs also matters. We have not measured those costs in this incident, but they belong in a business evaluation.
What self-hosted inference gives you control over
Inference is the processing that runs a model on an input and produces its output. With a suitable model running entirely on dedicated hardware you control, those inference requests do not consume a hosted model provider's token allowance.
You can reserve the machine for a particular workflow, set concurrency, prioritize urgent requests, and schedule batch work during quiet hours. You also choose when to change the model or serving software. These choices make it possible to build a capacity plan around your own workload.
This means deploying a separate model and workflow suited to your business. It does not mean moving the hosted Dot service onto a Mac or assuming a local model will match its capabilities. Test the quality of the actual task before moving it.
For the campaign example, local inference could handle suitable research summarization, classification, or drafting steps. The sending service remains a separate dependency. Its quotas, account controls, and delivery behavior still apply, as do limits on any external research tools. Moving inference does not remove those constraints.
Predictable capacity comes from measurement and reservation
A machine has finite memory and processing capacity. Longer inputs, longer outputs, different models, and more simultaneous users can change how quickly jobs finish.
For example, Ollama's documentation describes memory-dependent concurrency, queued requests, and overload errors when the queue becomes too large. Running locally gives the operator control over those settings; it does not make congestion disappear.
A useful capacity test uses representative business jobs and measures completed, acceptable results. Record how long requests wait, how long they take, how many fail, and how performance changes under simultaneous use. Serving tools such as vLLM expose queue, latency, and running-request metrics to help make bottlenecks visible.
Consider a deliberately simplified planning example. Assume a particular model and machine sustain 40 acceptable drafting jobs per hour for a defined input and output size. This is a hypothetical assumption, not a Looski benchmark. If you admit 30 jobs per hour, you leave 25% of that assumed throughput unused as headroom. A batch of 90 jobs then needs three hours of scheduled capacity, before human review or external sending delays.
That arithmetic is useful only while the workload and measured performance stay within the tested range. Larger jobs, retries, another workload sharing the machine, or a failure can change the result. The operator must monitor those conditions and adjust admission accordingly.
What would a capacity guarantee actually require?
Dedicated hardware lets you reserve resources for your firm. A promise about completed work needs more: a defined workload, an admission limit, tested performance, monitoring, and a recovery plan.
For an important daily workflow, specify the input-size range, maximum concurrent jobs, completion target, and what happens when demand exceeds the plan. Reserve its processing window and decide how unfinished jobs carry over. Record completed steps so recovery can avoid repeating actions such as sending an email. Keep spare capacity or a tested backup route if work must continue during a machine failure. Establish who responds when the service is unhealthy.
An owned computer can still lose power, fail, or encounter a software problem. Hardware, electricity, support, repairs, and replacement also remain costs. Predictability improves when those responsibilities are explicit; “no surprises” is too broad a promise.
Reserved capacity is also available in hosted services. For example, Amazon Bedrock offers Provisioned Throughput for supported models. A business should compare the available deployment options against its required quality, capacity, data boundary, and operating effort.
The case for capacity your business controls
Our earlier experience with Dots showed how quickly a connected assistant could help us get work started. This interruption adds another criterion: what happens when the workflow becomes busy enough to matter?
At Looski, the case for self-hosted inference is the ability to match a system to a firm's recurring work, measure its performance, and reserve resources for that work. Routine tasks can have explicit queue limits and priorities, with clear responsibility for keeping the service running.
Before depending on an AI workflow, ask what capacity is available, what consumes it, and what happens when it runs out. A business can plan around a known constraint. It needs a better operating answer than “check back in a bit.”
Frequently asked questions
What is Dot’s exact rate limit?
We have not established it. The message did not provide a numerical threshold, unit of usage, or reset time. Our interruption lasted several hours, but that observation does not establish a general cooldown or an emails-per-day allowance.
Does the message prove our campaign was considered abusive?
No. It identifies an abuse-prevention limit but does not explain the trigger. A usage threshold, burst control, or activity-specific check are possible explanations; the screenshot cannot distinguish them.
Does self-hosted inference guarantee uninterrupted AI work?
No. It lets you dedicate and manage computing resources. A completion or availability commitment needs a defined workload, tested capacity, controlled demand, and recovery arrangements. Hardware and software failures remain possible.
Will local inference remove outbound email limits?
No. Generating content locally and sending email are separate steps. The email provider and external research services retain their own quotas and account controls.