Economics and buying Guide
Self-Hosting LLMs in 2026: When Open-Weight Models on Your Own GPUs Pay
Open-weight models are now about four months behind the best closed models, and the GPUs to run them rent for a few dollars an hour. That makes self-hosting look cheap on a spreadsheet. The full cost depends on how busy the GPUs stay, what the cheapest host charges for the same weights, and who keeps the serving stack patched. This guide sets out when owning the model pays and when it only moves the bill.
For CTOs, heads of platform and data leaders deciding whether to run open-weight language models on their own or rented GPUs, use a hosted open model, or stay with a closed model API.
The short answer
Self-hosting an open-weight LLM pays when data must stay under your own control, when a narrow task runs well on a small model at very high volume, or when one shared GPU pool stays busy across many workloads. For everything else, an API or a hosted open model is usually cheaper, because idle GPUs cost many times more per token than busy ones.
Key takeaways
- The best open-weight models trail the best closed models by about four months on Epoch AI's capability index, and the top closed model led the top open one on Arena by 3.3% in March 2026.12
- In a 2026 field study, a developer on a self-hosted open model made 2.4 times as many fix commits as on a frontier closed model over the same length of time.3
- The price to beat is the same open model on a hosted API: gpt-oss-120b costs $0.15 per million input tokens and $0.60 per million output tokens on Together.4
- On identical H100s, measured cost ranged from $0.21 to $15.25 per million output tokens depending on load, so utilization decides whether owned GPUs are cheap.5
- The serving stack carries its own risk: a critical remote-code-execution flaw in vLLM in 2026, and about 175,000 Ollama servers found open to the internet.67
- Control of the logs is the lasting reason to self-host: from June 9, 2026, Anthropic keeps prompts to its most capable models for 30 days even under zero data retention.8
Self-hosting a large language model means running open-weight model files, such as DeepSeek, Kimi, Qwen, gpt-oss, Gemma or Llama, on GPUs you own or rent, with your own serving software, instead of sending prompts to a provider's API. In 2026 the case for it looks stronger than ever. Open models have closed most of the quality gap, GPU rental prices have fallen, and the serving software is mature and free. The case against is also stronger. Hosted providers sell the same open weights for cents per million tokens, prompt caching has cut the real cost of closed APIs, and the serving stack has a long security record. The useful question is which of your workloads pay for owning the model and which only move the bill.
- 4 monthsaverage lag of the best open-weight models behind the best closed models since January 20261
- 2.4xas many fix commits on a self-hosted open model as on a frontier closed model, in one developer's 28-day comparison3
- 36.3xhigher cost per token near idle than at full load, measured on the same H100 hardware5
How far behind the open models are
Every independent tracker still puts a closed model on top. Epoch AI found that since January 2026 the most capable open-weight models have lagged frontier closed models by an average of four months, a gap of 8 points on its capability index, "similar to the gap between GPT-5 and GPT-5.5".1 Mozilla's State of Open Source AI report measured the lag at about 4.4 months.9 Stanford's AI Index found the gap reopening: in March 2026 the top closed model led the top open model on Arena by 3.3%, up from 0.5% in August 2024.2 On Artificial Analysis's index, the highest-scoring open models in October 2026 are MiMo-V2.6-Pro and GLM-5.3, followed by Kimi K3.10
Four months sounds small, and on open-ended work it shows up as engineering time. In a 2026 field study by engineers at Pegatron in Taiwan, one developer spent two 28-day periods on the same production codebase, first with Claude Opus 4.7 and 4.8 through an API, then with GLM-5.1 and 5.2 running on the company's own Blackwell GPUs. Fix commits rose from 113 to 275, and the share of commits that were fixes rose from 46% to 75%.3 Government testing points the same way. NIST's CAISI found Kimi K2 Thinking below leading US models on agentic software and cyber tasks, "even older models such as GPT-5 and Opus 4".11
Narrow tasks are different, and that is where open weights earn their place. A fine-tuned Qwen2.5-0.5B, a model small enough to run on a laptop, reached a micro-F1 of 0.83 on relation extraction, against 0.69 for GPT-5.4 and 0.66 for Claude Sonnet 4.6 used without examples.12 A fine-tuned Phi-3.5 Mini matched GPT-4o on enterprise search relevance labeling at much higher throughput.13 The tasks where small open models win are classification, extraction, labeling and routing, at high volume, with a test set to prove it.
The open models worth hosting, and what their licenses allow
| Model | Size | What it takes to serve | License and catch |
|---|---|---|---|
| Kimi K3 | 2.8 trillion parameters, 104 billion active14 | About 1.4 TB of weights at 4 bits by our arithmetic, nearly all the memory of an eight-GPU B200 server before any cache | Modified MIT. A separate agreement if you sell it as an API with over $20 million revenue in 12 months, and "Kimi K3" shown in products above 100 million monthly users or $20 million monthly revenue15 |
| DeepSeek-V4-Pro | 1.6 trillion parameters, 49 billion active, 1 million token context16 | A full eight-GPU server or more | MIT. No usage triggers; provenance review advised (see below) |
| gpt-oss-120b | 117 billion parameters, 5.1 billion active | "A single 80GB GPU (like NVIDIA H100 or AMD MI300X)"17 | Apache 2.0 |
| Gemma 4 | Up to 31 billion parameters | One data-center GPU at the largest size | Apache 2.0 since April 2, 202618 |
| Mistral Medium 3.5 | 128 billion parameters | Two to four 80 GB GPUs, depending on precision | No rights at all if your company's monthly revenue exceeds $20 million19 |
| Llama 4 | Several sizes | One GPU to a full server | A separate license above 700 million monthly users and "Built with Llama" shown;20 multimodal rights withheld from companies based in the EU21 |
For a company using a model inside its own products, most of the open frontier carries no usage triggers. The traps are specific. Mistral Medium 3.5 withdraws every right from any company whose global monthly revenue exceeds $20 million, about $240 million a year, and that covers derivatives too.19 Kimi K3's revenue gate applies only to companies selling the model as a service, and internal use is exempt.15 Llama 4's acceptable use policy withholds rights to its multimodal models from any company with its principal place of business in the EU.21
Provenance is the harder question, because the strongest open models today are mostly Chinese-built. Mozilla counted eight open-weight models among the top ten on OpenRouter by token volume in August 2026, seven of them from Chinese labs.9 Self-hosting removes the data-flow risk behind the 2025 government bans on DeepSeek's app. It does not remove behavior trained into the weights. NIST's CAISI found agents built on DeepSeek R1-0528 were on average 12 times more likely than US frontier models to follow instructions planted to hijack them, and under a common jailbreak the model answered 94% of overtly malicious requests, against 8% for US reference models.22 Those tests cover 2025 models. For agents on any open weights, run your own red-team tests before giving the model tools.
The model file is a supply-chain risk of its own. By April 2025, Protect AI had scanned 4.47 million model versions on Hugging Face and flagged 352,000 unsafe or suspicious issues across 51,700 models.23 Load weights in the safetensors format from the publisher's official repository, pin the exact revision, and keep remote code execution switched off.
What self-hosting has to beat on price
Most business cases for self-hosting compare owned GPUs at full load with closed-API list prices. Both halves of that comparison are wrong. On the closed side, Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, and Claude Opus 5.5 costs $4 and $20,24 while OpenAI's GPT-6.1 Sol costs $2 and $10 and GPT-6 Astra $10 and $50.25 On the open side, the same weights you would host are already for sale. Together charges $0.15 and $0.60 for gpt-oss-120b and $2.70 and $13.50 for Kimi K3.4 DeepInfra charges $1.30 and $2.60 for DeepSeek-V4-Pro. It sells Llama 3.3 70B for $0.10 and $0.32, while Together charges $1.04 for the same model, so hosts of identical weights differ by up to ten times.264
GPU rental has fallen to a few dollars an hour. DeepInfra lists dedicated H100s at $2.20 per GPU-hour and B200s at $3.69.26 Lambda lists H100s at $4.29 and B200s at $6.99,27 RunPod lists B200s at $6.79,28 and AWS Capacity Blocks price an H100 at about $5.19 per GPU-hour in its main US regions.29 Two eight-GPU H100 servers at DeepInfra's rate, so that one can fail without an outage, cost about $25,700 a month. Add one engineer at the US median wage for software developers, $135,980 in May 2025,30 and the bill is about $37,000 a month before benefits, networking or a second person for on-call cover.
Against a closed frontier model, two rented servers break even only if the open model is good enough for the task and the cluster stays busy enough to produce several billion tokens a month. Against the same open model on a hosted API, they need about 141 billion tokens a month, more than 3,000 tokens a second from every GPU, every hour of the month. A 2026 study of H100 serving shows why that rarely happens. On identical hardware, effective cost ranged from $0.21 to $15.25 per million output tokens: 2.5 to 24 times more at the 1 to 10 requests a second typical of enterprise traffic, and up to 36.3 times more near idle.5 Hosted providers pool demand from many customers to stay near the cheap end. A single company rarely can.
Prompt caching moves the comparison further. On Claude, reading cached input costs a tenth of the normal input price, and 5% on Opus 5.5.24 In the field study above, a 99.3% cache-hit rate cut the closed API's cost by 88.6%, to an effective $0.57 per million tokens, below the $2.83 per million of the company's shared on-premises GPUs.3 On-premises still saved 40.1% of total cost there, because the GPUs were a shared allocation and labor was priced at $35 an hour. A dedicated reservation of the same GPUs cost 43.8% more than the cached API.3
Time works against hardware bought at today's prices. a16z found that for a model of equivalent performance, cost falls "by 10x every year".31 Epoch AI found a median fall of 50 times a year across benchmarks, rising to 200 times a year on data since 2024.32 A Carnegie Mellon cost model that counted hardware only found break-even "within a few months for small models, 2 years for medium models and 5 years for larger models", viable mainly at very high volume or under strict data residency mandates.33 Plan GPU purchases over three years at most, and expect to swap the model on them more than once.
Running the stack: engines, upgrades and silent drift
The serving layer has consolidated. vLLM became a foundation-hosted project of the PyTorch Foundation in May 202534 and shipped version 0.31.0 in October 2026.35 LinkedIn runs it for more than 50 generative AI use cases on thousands of hosts,36 and Roblox reported an almost 2x improvement in latency and throughput after moving to it, serving about 4 billion tokens a week.37 Hugging Face put its own server, TGI, into maintenance mode and archived the repository on March 21, 2026.38 On Kubernetes, llm-d, which routes requests to the GPU that already holds a prompt's cache, became a CNCF Sandbox project in March 2026.39
Compressing a model to save memory has a measured cost. Across five models and inputs longer than 64,000 tokens, 8-bit quantization lost about 0.8% accuracy on average, while 4-bit methods lost up to 59%, and the damage grew when the input was not in English.40 In another study, a 1.7% average drop in Japanese on automatic tests corresponded to a 16.0% drop judged by people on realistic prompts.41 Use 8-bit by default and test 4-bit on your own long and non-English inputs before trusting it.
The most expensive failure is silent: the same model giving different answers depending on who serves it and how. When Artificial Analysis ran gpt-oss-120b on a math benchmark across providers, most scored 93.3%, Azure scored 80.0% and one heavily compressed deployment scored 36.7%, and most of the top scorers were running the latest vLLM.42 Moonshot's own verifier found Kimi K2 Thinking's tool-call schema accuracy ranging from 100% at its own API down to 83% at other providers.43 The vLLM team found that fewer than 20% of 1,200 potential Kimi K2 tool calls parsed correctly before it fixed template and parser bugs, after which 99.925% of requests succeeded.44 A self-hosting team inherits this problem. Every engine upgrade, template change or quantization switch needs the same evaluation run before release, including tool calls, long inputs and every language you serve.
Hardware fails more often than teams expect. A study of 11.7 million GPU hours at NCSA found H100 memory had 3.2 times lower mean time between errors than A100 memory, and concluded that 5% spare capacity is needed to absorb GPU failures.45 Budget for a spare node, or for the share of traffic you can afford to drop while a server reloads a model.
The security record of self-hosted serving
| Date | What happened | The control that stops it |
|---|---|---|
| April 2025 | Protect AI's scans of Hugging Face flagged 352,000 unsafe or suspicious issues across 51,700 models23 | Load safetensors from official repositories, pinned to a revision |
| November 2025 | Oligo found the same unsafe pattern, Python pickle data read over ZeroMQ sockets, copied across vLLM, SGLang and NVIDIA's TensorRT-LLM46 | Keep internal engine ports on a private network; patch on the engine's release cycle |
| November 2025 | ShadowRay 2.0 exploited a 9.8-rated missing-authentication flaw in Ray (CVE-2023-48022) to build a cryptomining botnet; more than 230,500 Ray servers were publicly reachable47 | Never expose cluster dashboards or job APIs to the internet |
| January 2026 | SentinelOne and Censys found about 175,000 Ollama servers open to the internet in 130 countries, nearly half with tool calling enabled7 | Treat developer tools as developer tools; put production serving behind authentication |
| February 2026 | CVE-2026-22778: a chain in vLLM's video processing allowed remote code execution through a crafted video URL, rated critical at 9.8, fixed in version 0.14.16 | Switch off input types you do not use and track the engine's advisories48 |
Where the data goes with each option
Between a public API and your own GPUs sits a range of options that answer most residency needs. What none of them offers is a provider with no retention or review path at all, and 2026 made that plainer.
| Option | What the provider says |
|---|---|
| OpenAI API | Abuse-monitoring logs kept "for up to 30 days"; zero data retention needs OpenAI's approval; data residency available in regions including India49 |
| Anthropic, most capable models | From June 9, 2026, prompts and outputs of covered models are kept for 30 days "on every platform where these models are offered", including zero-retention workspaces and AWS, Google Cloud and Microsoft clouds8 |
| Models sold by Azure in Microsoft Foundry | Prompts and outputs "are NOT available to OpenAI"; processed in the chosen geography unless you use a Global or DataZone deployment50 |
| Amazon Bedrock | Models run in deployment accounts owned by AWS, and "model providers don't have any access to those accounts";51 you can import your own gpt-oss weights52 |
| Hosted open models on Together | Unless zero data retention is enabled, Together "stores the prompts you send and the responses models return, and may use them for product improvements"53 |
| Disconnected and air-gapped | Gemini on Google Distributed Cloud air-gapped is generally available;54 large models on Foundry Local are available to qualified customers55 |
| Your own GPUs | Retention, review and logs are whatever you decide |
In Europe, the pressure comes from US law reaching US providers. Asked at a French Senate hearing in June 2025 whether data held for French public buyers would never be passed to US authorities without French consent, Microsoft France's general counsel said he could not guarantee it under oath.56 India's DPDP Rules allow transfers abroad subject to requirements the central government may set on making data available to a foreign state.57 Open weights on infrastructure you control are the one setup that no provider's terms or foreign order can change.
The EU AI Act barely changes the decision. Under the Commission's guidelines, a company that modifies a general-purpose model becomes its provider only when the modification uses more than a third of the original model's training compute,58 far beyond ordinary fine-tuning. The duties that follow from what the system does apply whether you call an API or run your own GPUs.
When self-hosting pays
| Workload | Default in 2026 | Self-host when |
|---|---|---|
| Agentic coding, complex reasoning, long multi-step tool use | Closed model API with caching and batch discounts | Rarely; the quality gap comes back as rework |
| Narrow, high-volume tasks: classification, extraction, routing, scoring | Small fine-tuned open model on a hosted endpoint | Volume is in the billions of tokens a month and steady, or the data cannot leave your control |
| Regulated data that needs in-region processing | Regional deployment of a closed model, or an open model on a dedicated regional endpoint | Provider retention or review is unacceptable to your regulator or board |
| Classified, air-gapped or sovereign-only data | Open weights on your own GPUs, or an air-gapped offering | Always; this is the core case |
| Many internal workloads and an existing platform team | One shared GPU pool behind a gateway | Combined traffic keeps the pool busy most of the day |
| Pilots and spiky or low volume | Any API | Never; idle GPUs are the most expensive way to buy tokens |
The pattern that holds up is a gateway that routes each request by data class and task difficulty. Uber's GenAI Gateway gives its developers one interface to models from OpenAI and Vertex AI and to models Uber hosts itself, and runs a PII redactor that "anonymizes sensitive information within requests before forwarding them to third-party vendors".59 Hard reasoning goes to the best closed model, narrow volume goes to a small open model, and data that must stay inside goes to GPUs you control.
Closed model API
Best quality, least operations
- Frontier quality on hard tasks
- Caching and batch cut real cost sharply
- Provider retention and review paths apply
- Prices and models change on the provider's schedule
Hosted open model
Same weights, someone else's GPUs
- Cents per million tokens for mid-size models
- Swap hosts without changing the model
- Check retention defaults and serving quality
- Dedicated regional endpoints for residency
Open model on your own GPUs
Full control, full responsibility
- Logs and retention entirely yours
- Cheap only while GPUs stay busy
- Patching, evaluation and spare capacity are yours
- Hardware outlives several model generations
How to decide in six weeks
- Classify the data Sort each workload by what its prompts contain and which rules apply, and mark the ones that may never leave your control.
- Measure the volume Record monthly tokens, the input-to-output ratio, peak requests per second and how much of the input repeats and could be cached.
- Build the test set Collect a few hundred real cases per workload, including tool calls, long inputs and every language you serve, with agreed pass marks.
- Test open models hosted first Run the candidate open models on a hosted endpoint against the test set before buying or renting any GPUs.
- Price at measured load Cost the self-hosted option at the utilization your traffic actually produces, with spare capacity, staff and on-call included, against the cheapest host of the same model.
- Move only what passes Self-host the workloads that pass on quality and either pay on cost or must stay inside; leave the rest on APIs behind the same gateway.
- Gate every upgrade Rerun the test set before any engine, model, template or quantization change reaches production.
Questions before buying or renting GPUs for LLMs
- Which workloads must stay under our control, and which only prefer to?
- What does the cheapest host charge for the exact model we plan to run?
- What utilization will our real traffic produce, hour by hour?
- Does the open model pass our own test set, including tool calls and long inputs?
- Does its license allow our use at our revenue and user numbers?
- Who patches the serving stack, and how fast after an advisory?
- How do we keep serving when a GPU or a whole server fails?
Running your own model in 2026 is mostly a decision about control. The cost case survives only for steady, high-volume work on models small enough to keep GPUs busy, because the falling price of hosted tokens, prompt caching and idle hardware all work against it. The control case has grown stronger as provider retention terms gained exceptions. Self-host the data that has to stay inside, rent capability for everything else, and keep the switch between them in one gateway so the line can move as prices and models do.
This is how we approach model hosting in our AI development work: classify workloads by data and difficulty, test open models on a hosted endpoint before any hardware decision, cost self-hosting at measured load, and put every model behind one gateway with an evaluation gate on each upgrade.
Questions leaders ask
Is it cheaper to self-host an LLM than to use an API?
Only when the GPUs stay busy. On identical H100s, measured cost ranged from $0.21 to $15.25 per million output tokens depending on load. At typical enterprise traffic, hosted providers that pool demand from many customers usually charge less than a company pays for its own idle capacity.
How good are open-weight models compared with GPT and Claude?
The best open-weight models trail the best closed models by about four months of progress on Epoch AI's capability index. On narrow tasks such as classification and extraction, small fine-tuned open models can match or beat frontier models. On agentic coding and complex reasoning the gap still shows up as extra rework.
What hardware do I need to self-host an LLM?
It depends on the model. gpt-oss-120b fits a single 80 GB GPU such as an H100. The largest open models, such as Kimi K3 at 2.8 trillion parameters, need one or more full eight-GPU servers. Plan a spare server or spare capacity, since GPUs fail and models take minutes to reload.
Is it safe to use Chinese open-weight models like DeepSeek or Kimi?
Running the weights on your own servers keeps data away from the developer's services, which is what 2025 government bans on DeepSeek's app targeted. Behavior trained into the weights remains, and NIST found DeepSeek R1-0528 far easier to hijack and jailbreak than US models. Test any open model yourself before giving it tools or sensitive tasks.
Which software is used to serve open LLMs in production?
vLLM is the most widely used open serving engine and is hosted by the PyTorch Foundation, with SGLang and NVIDIA's TensorRT-LLM as the main alternatives. Ollama and llama.cpp suit development and edge devices. Hugging Face archived its TGI server in March 2026.
Does self-hosting help with GDPR, DPDP or data residency?
It gives you full control of where prompts go and how long they are kept, which no provider contract fully matches. Regional deployments of closed models and dedicated open-model endpoints meet many residency needs at lower cost, so self-hosting matters most where provider retention or foreign legal access is unacceptable.
Does fine-tuning an open model make us a provider under the EU AI Act?
Usually not. The European Commission's guidelines treat a company that modifies a general-purpose model as its provider only when the modification uses more than a third of the original model's training compute, far beyond ordinary fine-tuning.
Sources
- Open models lag state-of-the-art closed models by 4 monthsEpoch AI, May 2026
- AI Index Report 2026: technical performanceStanford HAI, 2026
- Inference economics of enterprise coding agents: a case study of cloud vs. on-premise LLMsPeng et al., arXiv 2607.13080, July 2026
- PricingTogether AI
- Beyond per-token pricing: a concurrency-aware methodology for LLM infrastructure cost estimationPatil, arXiv 2606.11690, June 2026
- vLLM: remote code execution via a crafted video URL (CVE-2026-22778)GitHub Advisory Database, February 2026
- Researchers find 175,000 publicly exposed Ollama AI servers across 130 countriesThe Hacker News, January 2026
- Data retention practices for covered modelsAnthropic Privacy Center
- State of Open Source AIMozilla, September 2026
- Open weights models: intelligence indexArtificial Analysis
- CAISI evaluation of Kimi K2 ThinkingNIST, December 2025
- Sub-billion, super-frontier: small language models rival zero-shot frontier LLMsarXiv 2606.22606, June 2026
- Fine-tuning small language models as efficient enterprise search relevance labelersarXiv 2601.03211, January 2026
- Kimi K3 model cardMoonshot AI, Hugging Face
- Kimi K3 licenseMoonshot AI, Hugging Face
- DeepSeek-V4-Pro model cardDeepSeek, Hugging Face
- gpt-oss-120b model cardOpenAI, Hugging Face
- Gemma 4: expanding the Gemmaverse with Apache 2.0Google Open Source Blog, 2026
- Mistral Medium 3.5 licenseMistral AI, Hugging Face
- Llama 4 Community License AgreementMeta
- Llama 4 Acceptable Use PolicyMeta
- CAISI evaluation of DeepSeek AI models finds shortcomings and risksNIST, September 2025
- 4M models scanned: Protect AI and Hugging Face six months inHugging Face, 2025
- Claude API pricingAnthropic
- API pricingOpenAI
- PricingDeepInfra
- PricingLambda
- PricingRunPod
- Amazon EC2 Capacity Blocks for ML pricingAWS
- Software developers, quality assurance analysts and testers: payUS Bureau of Labor Statistics, May 2025 data
- Welcome to LLMflation: LLM inference cost is going down fasta16z, 2024
- LLM inference prices have fallen rapidly but unequally across tasksEpoch AI
- A cost-benefit analysis of on-premise large language model deploymentCarnegie Mellon University, arXiv 2509.18101, 2025
- PyTorch Foundation expands to an umbrella foundation and welcomes vLLM and DeepSpeedPyTorch Foundation, May 2025
- vLLM releasesvLLM project, GitHub
- How we leveraged vLLM to power our GenAI applicationsLinkedIn Engineering
- Running AI inference at scale in the hybrid cloudRoblox, September 2024
- text-generation-inference repositoryHugging Face, GitHub
- llm-d is officially a CNCF Sandbox projectGoogle Cloud, 2026
- Does quantization affect models' performance on long-context tasks?arXiv 2505.20276, EMNLP 2025
- How does quantization affect multilingual LLMs?arXiv 2407.03211, 2024
- Open weight LLMs exhibit inconsistent performance across providersSimon Willison, August 2025
- K2 Vendor VerifierMoonshot AI, GitHub
- Chasing 100% accuracy: a deep dive into debugging Kimi K2's tool calling on vLLMvLLM blog, October 2025
- Story of two GPUs: characterizing the resilience of Hopper H100 and Ampere A100 GPUsNCSA, arXiv 2503.11901, 2025
- ShadowMQ: how code reuse spread critical vulnerabilities across the AI ecosystemOligo Security, November 13, 2025
- ShadowRay 2.0 exploits unpatched Ray flawThe Hacker News, November 2025
- Security advisoriesvLLM project, GitHub
- Your dataOpenAI API documentation
- Data, privacy and security for Foundry Models sold by AzureMicrosoft Learn
- Data protection in Amazon BedrockAWS
- Amazon Bedrock Custom Model Import adds OpenAI gpt-oss modelsAWS, November 2025
- Zero data retentionTogether AI documentation
- Google Cloud makes Gemini everywhere vision a realityGoogle Cloud, August 2025
- Microsoft Sovereign Cloud adds governance, productivity and support for large AI models running disconnectedMicrosoft, February 2026
- Not sovereign: Microsoft cannot guarantee the security of EU dataheise online, June 2025
- DPDP Rules 2025, Rule 15: transfer of personal data outside Indiadpdpa.com
- Guidelines on the scope of the obligations for general-purpose AI modelsEuropean Commission, July 2025
- Navigating the LLM landscape: Uber's innovation with GenAI GatewayUber Engineering
Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 6, 2026. No client data appears in our insights.