
An email from Google landed in our inbox with a familiar subject line: action required. Two Gemini models are being retired. Applications using them need to move.
For Google, it is a model lifecycle update. For a business that has built a working process around one of those models, it is an unplanned project with a deadline.
That is one of the costs of building on AI you access through someone else's service: you can choose when to start using a model, but the provider can decide when you must stop.
What Google's notice says
The Google Cloud notice we received lists these retirement dates for Gemini Enterprise:
| Model | Retirement date in the notice | Stated consequence |
|---|---|---|
| Gemini 3.6 Flash | November 19, 2026 | Requests to the retired endpoint will return HTTP 404 errors |
| Gemini 3.7 Flash | January 28, 2027 | Requests to the retired endpoint will return HTTP 404 errors |
The notice recommends testing and migrating to Gemini 3.8 Flash or Gemini 3.5 Flash. Customers with Provisioned Throughput must also move that capacity to a supported model.
These dates come from the notice received by Looski. When checked on October 5, 2026, Google's public lifecycle table still showed no announced retirement date for either model. The table identifies both as short-term availability models. Readers should check their own project notices and current documentation before scheduling a migration. The email explicitly limits its scope to Gemini Enterprise; Gemini API and AI Studio schedules can differ.
The operational issue is clear even before the public table catches up: a workflow can still be useful to its owner when its provider schedules the underlying model for retirement.
The work goes beyond changing a model name
Changing the endpoint may take minutes. Establishing that the replacement still does the job can take much longer.
Consider a business using AI to read incoming requests, extract a few fields, and prepare a reply for staff approval. Its prompts, examples, and checks have been adjusted around the existing model. A replacement may interpret an ambiguous request differently, produce a different structure, or require more staff corrections. Those are possibilities to test, not failures we have measured in these Gemini replacements.
Google's migration guide recommends repeating evaluations, checking individual components in complex workflows, and testing throughput before rollout. It also allows for prompt adjustments when results change. That is sensible engineering, and it takes time.
A successful migration can prevent an outage and still disrupt the business. Someone has to pause other work, compare outputs, investigate differences, and release the change. The API bill does not capture all of that effort.
For teams using several models across several workflows, retirement notices become a recurring operating responsibility.
A model version is part of your operating process
Businesses spend time getting a process right: the information it needs, the decisions it can support, the format it returns, and the point where a person reviews it. A model becomes one component of that process.
Access through a hosted API does not give the business control over how long that component remains available. Pinning a model identifier can help specify what you call today. It cannot keep an endpoint alive after the provider retires it.
The business question is therefore practical: if this model disappeared, how much of our process would we need to validate again, and how quickly could we do it?
Routers can soften the disruption
A router such as OpenRouter or Vercel AI Gateway gives your application one integration through which it can reach multiple models and providers. That can reduce the plumbing involved in a migration and help keep requests moving when an upstream service fails.
There are two distinct protections:
- Provider fallback tries another provider serving the same model, where one is available and allowed by your routing settings. This can help with an outage or capacity problem at one provider.
- Model fallback tries a different model you have selected. This is the more relevant protection when the original model is retired across its available providers.
OpenRouter supports an ordered list of fallback models, which it tries when a preceding model returns an error. Vercel AI Gateway also supports model fallbacks: it applies provider routing for each model, then moves to the next configured model if those providers fail.
For the incoming-request workflow, a team could validate a second model in advance and configure it as a fallback. If the primary becomes unavailable, the router can attempt that alternative through the same integration. A planned retirement should still trigger a deliberate switch before the deadline; fallback is useful protection while the change is being completed.
These services mitigate the dependency. They cannot keep a retired model alive or guarantee that a replacement behaves the same way. Simply adopting a gateway does not select and validate the right backup for your business. Test the fallback's output format, tool support, context limits, cost, and response time. Restrict eligible providers to those that meet your data-handling requirements, and record which model actually served each request.
The router itself is another service dependency, but for teams staying with hosted AI, an evaluated fallback behind a shared interface is a practical way to reduce disruption.
What owning your model should mean
For most businesses, training a foundation model from scratch is unnecessary. The useful form of ownership is operational control: retaining downloadable model weights under a suitable license and running them on infrastructure you control.
With that setup, a provider closing its hosted endpoint does not by itself shut down your installed copy. You can keep a validated version running while you evaluate a replacement on your own schedule.
That requires more than downloading a file. Keep the compatible tokenizer, runtime, configuration, and any adapters alongside the weights. Preserve the examples and checks that establish whether the workflow works. Test that you can restore the deployment.
This does not mean you own the model's intellectual property, that every open model has unrestricted terms, or that a local model will match the hosted service on your tasks. License, capability, memory, and throughput all need assessment. A downloaded model still used exclusively through somebody else's API leaves that service dependency in place.
Self-hosting also makes capacity, updates, monitoring, and recovery your responsibility. It buys control over one important dependency while adding operating work of its own.
Build continuity into the workflow
Our recommendation is to start with the processes where an interruption would hurt most.
- Know what depends on each model. Record the model identifier, service, workflow owner, and any announced retirement date.
- Keep a small, representative evaluation set. Include ordinary work, ambiguous inputs, and cases where an incorrect answer would create substantial rework. Compare accuracy, staff correction time, latency, and cost.
- Validate an alternative before you need it. That might be another hosted model, a self-hosted model, or a manual process. Check the whole workflow and its available capacity.
- Make upgrades deliberate. Test the replacement before moving production traffic. Keep a rollback path while the old endpoint remains available, and a continuity plan for after it closes.
Hosted models can be a good choice when their capabilities and convenience justify the dependency. For stable, repeatable work that a self-hosted model performs well, control over the upgrade schedule can be a substantial benefit.
The next model may be better. Your business still needs the freedom to decide when it is ready to change.
Frequently asked questions
When are Gemini 3.6 Flash and Gemini 3.7 Flash retiring?
The Google Cloud notice received by Looski gives November 19, 2026 for Gemini 3.6 Flash and January 28, 2027 for Gemini 3.7 Flash on Gemini Enterprise. The public lifecycle table checked on October 5 had not yet reflected those dates. Confirm your project notice and current documentation; AI Studio and Gemini API dates can differ.
Does changing the model identifier complete an AI migration?
It changes which model receives requests, but it does not establish that the replacement meets your workflow requirements. Compare outputs on representative inputs and check integrations, latency, throughput, staff correction time, and cost before rollout.
Does self-hosting eliminate AI service disruption?
No. A retained model and reproducible deployment can remove dependence on a provider continuing to serve that endpoint. Hardware failures, software changes, capacity limits, and operating mistakes remain your responsibility. You still need monitoring and recovery procedures.
Can OpenRouter or Vercel AI Gateway prevent disruption when a model retires?
They can reduce disruption by giving your application a shared integration and configurable fallback models. Provider fallback can try another host for the same model; model fallback can try a different supported model. Neither preserves a retired model or guarantees equivalent results. Validate alternatives, configure routing restrictions, and migrate deliberately before the retirement deadline.
Sources & further reading
- Google Cloud: model versions and lifecycle (checked October 5, 2026; email dates not yet reflected)
- Google Cloud: migrate to the latest Gemini models
- Google AI for Developers: Gemini API deprecations (separate service schedule)
- OpenRouter: automatic failover between models
- Vercel AI Gateway: model fallbacks and provider routing
- OpenRouter: provider routing and restrictions