Kubermatic CEO Scheele: Run AI models like software
Switching models can disrupt running AI processes. Platform teams need clear responsibility for updates and ongoing operations.
In some companies, AI teams fund their own platform, while the experience needed to run it over the long term sits with other teams. Kubermatic CEO Sebastian Scheele therefore argues for involving existing platform teams in operating AI models.
Key takeaways
- Models have a lifecycle. Sebastian Scheele calls for managed releases and the controlled replacement of old versions.
- Operating models requires clear responsibilities. Existing platform teams can bring their experience with containers to AI infrastructure.
- Utilization follows the workflow. Deferrable batch tasks can use spare capacity at night. The value of the result remains decisive.
Related:MLOps in the Cloud: Reliably Deploying Machine Learning Models into Production / Small Models Devour Large GPU Budgets through Preallocation

Operational work starts with the first model
Sebastian Scheele sees local AI as a realistic option for companies. Models are becoming more capable and cheaper to run, the CEO and co-founder of Kubermatic says in an interview at ContainerDays in Hamburg. His central requirement is: “I need to operate this like software.” He is focusing on the work that comes after installation.
In the process he describes, a team buys powerful hardware, installs a large model and develops processes around it. Those applications must continue to work when a new version appears. Switching models therefore also affects the teams that use the model’s answers in their workflows. The operator needs an overview of these dependencies.
“I really need release management and deployment management for these models.” For Scheele, this includes running several models in parallel, rolling out new versions and retiring old versions. He gives the example of a short transition period for a team still using an earlier version. The length of that transition period depends on the applications affected.
// Definition
What is release management for AI models? It means managing the operation of models after installation. This includes running several models in parallel, rolling out new versions and retiring old versions with a transition period for the teams still using them.
AI budgets cannot replace operational experience
In Scheele’s observation, AI or development teams have often received the budget for a platform. Experience in running it over the long term often sits elsewhere, however. This can create additional environments whose upkeep has to be organized separately from the existing infrastructure.
Scheele argues for involving the existing platform teams. They already run containers and can assess how AI workloads could be accommodated in their environment. The business units and development teams would then use the models provided, while the teams responsible for operations would manage the infrastructure. “It is still just software that needs to be operated.”
This division of responsibilities also makes the upkeep of the other components visible. Scheele describes the rapid pace of software changes in AI: some older versions stop receiving development work or patches after only a few months. For operations, what matters is which components are actually in use and how long they receive maintenance.
Night-time runs can use spare GPU capacity
When discussing utilization, Scheele connects infrastructure with business processes. Load increases as developers work with their environments during the day. It may be possible to shift additional tasks that do not need to be completed immediately. He cites batch processing, meaning grouped runs without ongoing user interaction, as an example of work that can take place at night.
This decision requires the operator to know the teams’ needs. Some results are required daily, others weekly or monthly. An analysis that only needs to be ready the next morning has different timing requirements from an interactive request. This distinction helps make more consistent use of the available hardware.
Scheele’s argument concerns planning: expensive hardware incurs costs even when it is idle. Whether an additional night-time run makes economic sense also depends on what it delivers.
The value must justify operating the model
Scheele therefore calls for comparing token consumption with the value of the result. A large number of generated answers is no more a measure of success than a constantly busy GPU. The question is whether a run produces the desired output or whether rephrasing the task would help.
For platform teams, this amounts to a connected set of operational responsibilities: they need to know which teams use a model, when those teams’ work requires computing power and how version changes can remain possible. A centrally managed offering can bring these decisions together. Any capacity savings must be demonstrated in the specific operating environment.
Frequently Asked Questions
What does operation after installation mean?
Models and their software environment need managed updates, clear responsibilities and a controlled transition away from old versions. Scheele describes this work as Day-2 operations.
Which AI tasks are suitable for night-time runs?
Deferrable batch tasks with enough time before the result is needed. They cannot replace interactive requests.
What does the number of tokens tell us about value?
It describes the volume of text a model processes. Whether a hardware run makes economic sense also depends on whether its result fulfils the required task.
Tobias Massow conducted the interview on 3 September 2026 at the Kubermatic stand at ContainerDays in Hamburg.
Editor's Picks
cloudmagazinOpenAI Pulls the Plug on Cursor: Your Own Key as Plan BcloudmagazinDownloadable Doesn’t Mean DeployablecloudmagazinAWS lets an AI agent join incident investigationsMore from the MBF Media Network
MyBusinessFutureWhen a German AI Model Truly Pays OffDigital ChiefsFour Stumbling Blocks: Why AI Projects Fail to Transition to Regular OperationsSecurityTodaySASE Sees HTTPS to LLM – Not the IntentImage source: AI-generated (September 2026), portrait: Kubermatic
Translated from the German original using artificial intelligence. The German version is authoritative.

