Evaluating Model Deployment Options Based on Token Cost & Sovereignty
As more customers consider deploying agents in production, the top 3 topics of discussion are:
agent accuracy,
token cost and
data sovereignty
As a result, there are newer deployment options being discussed: hyperscaler vs. on-premises, frontier vs. open-weight, etc. These deployment options can be evaluated on 4 different dimensions:
Agent Accuracy: Frontier models have the best accuracy on the most number of use cases, whereas open-weight models may shine in one set of use cases and may be poor for other use cases. For our computer use scenarios, open-weight models are at least 3 months, perhaps 6 months behind frontier models in terms of capabilities.
Token Cost: OpEx or CapEx. Variable OpEx token cost is a serious cause for concern for customers. Deploying open-weight models on-premises is one way to convert variable OpEx into predictable CapEx.
Data Sovereignty: Much has been written about the key IP for customers being their proprietary data. Customers are reluctant to send data directly to the frontier models because they are worried that their data will be used to train the models and eventually disrupt their own businesses. Customers are worried about sending data to the hyperscalers even the hyperscalers have committed to not sharing the data with the frontier model providers. Again, on-premises deployments are the best options to control data sovereignty.
Operational Complexity: These models are incredibly complex to deploy and operate. Not surprisingly the operational complexity of consuming these models from a cloud service is low, whereas running the models in an on-premises data center would be more complex.
We have summarized the current state of deployment options in the table below: