PromptArmor
Blogs

Vendors adding Fireworks? Kimi, DeepSeek and GLM are on the way

We've spotted international model use by third-party vendors early, by analyzing the changes to their AI inference providers.

Inference hosts
ModelsChinaChinaChinaFrance
Your vendors
GleanHarveyZendesk????

Which of your vendors are adopting new inference hosts and international models?

Vendors are turning to international AI models

Vendors with AI features are increasingly turning to international and open-weight models due to the cost difference compared to models from leading U.S. labs, and because these international models are becoming increasingly competitive on benchmarks.

Recently, Kimi K3 and DeepSeek V4 Flash 0731 have made the news, offering competitive performance at a steep discount compared to models like Anthropic's Fable and OpenAI's GPT Sol.

International vs U.S. model token costs

InternationalU.S.
ChinaDeepSeek V4 Flash
In$0.14
Out$0.28
ChinaDeepSeek V4 Pro
In$0.435
Out$0.87
ChinaZ.ai GLM 5.2
In$1.40
Out$4.40
FranceMistral Medium 3.5
In$1.50
Out$7.50
United StatesOpenAI GPT-5.6 Terra
In$2.00
Out$12.00
ChinaMoonshot Kimi K3
In$3.00
Out$15.00
United StatesAnthropic Sonnet 5
In$3.00
Out$15.00
United StatesAnthropic Opus 5
In$5.00
Out$25.00
United StatesOpenAI GPT-5.6 Sol
In$5.00
Out$30.00
United StatesAnthropic Fable 5
In$10.00
Out$50.00

USD per million tokens. Provider list prices, read August 9, 2026. Sources: DeepSeek, Z.ai, Mistral, Moonshot, OpenAI, Anthropic.

Furthermore, open-weight models do not impose API-level restrictions on security and biology work, unlike leading U.S. model providers.

While organizations are not all aligned on the potential for risks such as model backdoors, most organizations are in agreement that data processing for AI inference must remain within the U.S.

Because of this, vendors that want to add international or open-weight models turn to inference providers whose job it is to run those models within U.S. data residency.

Inference provider changes predict model changes

Across the vendors we monitor for our customers, we see AI subprocessor changes frequently. From new features that come with new models to updating features with the latest models, subprocessor and model changes are common occurrences. Throughout these changes, there have been some notable trends.

One sequence we have seen play out again and again is that shortly after a vendor adds a US-based inference provider like Fireworks, Baseten, Together, Groq, or Cerebras, they begin rolling out open-weight or international models.

Below are samples of vendors we have seen take this path, and vendors who have recently added an inference provider but are yet to announce their more-than-likely function: supporting the adoption of international or open models.

Recent examples from vendors

International models following new inference subprocessor
Glean
Subprocessor addedJun 16, 2026
Fireworks
Model addedJul 16, 2026
ChinaZ.ai GLM 5.2
Harvey
Subprocessor addedJul 3, 2026
BasetenFireworks
Model addedAug 5, 2026
FranceMistral Medium 3.5
Vercel
Subprocessor addedJan 29, 2026
CerebrasBaseten
Model addedApr 8, 2026
ChinaZ.ai GLM 5.1
GitHub Copilot
Subprocessor addedAug 29, 2025
Fireworks
Model addedAug 6, 2026
ChinaMoonshot Kimi K3
New inference subprocessor, new models coming soon
Sierra
Subprocessor addedJun 3, 2026
BasetenFireworksTogetherModal
Model added
Coming soon
Zendesk
Subprocessor addedJul 17, 2026
GroqFireworksBaseten
Model added
Coming soon
Ada
Subprocessor addedJun 2026
GroqBaseten
Model added
Coming soon
Deepgram
Subprocessor addedMay 2026
Baseten
Model added
Coming soon

More predictive alerts for AI in vendors

Beyond the identification of international and open-weight model use, we have observed a number of trends that can be extrapolated to inform decisions about AI in vendors. Here are a few:

  • Cyber-capable model adoption. Anthropic added a “Covered Models” section to its Service Specific Terms on June 8, 2026, one day before launching Claude Fable 5 as a covered model on June 9, with supplemental terms that permit human safety review and explicitly supersede zero data retention commitments for their models that have advanced cyber capabilities. The upstream data processing requirements became a signal for the adoption of these models. For example, Glean introduced a Limited Retention Addendum on June 10, two days after Anthropic’s terms change, and added Claude Fable 5 within a month.
  • AI feature expansion. Across industries, adoption of new tools and AI capabilities usually occurs in groups. For example, Skills first spread across coding agents, and then shortly after they spread across knowledge work. Once one major player in an industry adopts a trending AI capability, competitors are quick to adopt - so by tracking the first entrant in an industry, we can predict when other industry players will adopt. This is playing out right now with Legal AI; Legora introduced Skills on June 3, 2026; since then, Harvey and other competitors have begun rolling out Skills. This pattern has played out multiple times now, including with Memory, Web Search, MCP, and recently, 'Coworkers' and always-on-agents.
  • Agentic platform access via nth party apps. As applications are increasingly accessed via AI connectors, vendors have begun to shift their posture to reflect the utilization of their services via third-party AI services, like Claude and ChatGPT. Vendors have begun to release terms changes that indicate when and how users are liable for access to their services via third-party AI, such as MCP servers and connectors. In addition to legal changes, a common tendency is to have a staggered rollout of ChatGPT and Claude connectors: if we see a vendor become accessible via ChatGPT or Claude, we can infer that it will soon be accessible through the other.
  • Agentic scope expansion. Vendors love to describe their features as 'agents' and 'agentic', but it is not uncommon to find 'agents' with limited practical capabilities beyond a normal chatbot. In the process of reviewing these agents, we've come across a pattern: from the first time an 'agentic' system surfaces in marketing, it takes about six months to go from a 'chatbot with a specialized prompt' to true agentic, autonomous, or always-on capabilities (MCP, web browsing, recurring workflows, code execution, etc.). Monitoring 'agents' for this transition has become a priority for many organizations that have approved these 'agents' on the basis of the limited toolset they originally presented.
  • Connector and integration releases. Recently, we released research that found that Claude and ChatGPT connectors change on average every 9 minutes. But the big vendors are not the only ones with tumultuous connector ecosystems. Across vendors of varying sizes with AI features, we’ve seen that the overhead to create one connector vastly lowers the prerequisites for future connectors. Shortly after we detect a vendor adding its first connector, they often enter a period of rapid growth in which dozens or hundreds of connector options are added, before the increase tapers back down and lands at a steady rate.

By monitoring vendors across all major industries and serving customers with priorities from security to TPRM to GRC to legal, and more, we gain access to bigger-picture insights about the ecosystem of AI in vendors. This enables us to convert those bigger-picture insights into predictive alerts, helping our customers stay ahead of the curve as their vendors expand and change their AI posture.

Get predictive alerts for AI in your vendors

PromptArmor Threat Intelligence

Is your organization protected from AI in vendors?

PromptArmor continuously monitors across your portfolio of third party AI in vendors, skills, plugins, connectors, MCP servers, models and more.

We detect vulnerabilities and changes like this, surfacing risk before it becomes an incident.

Learn more