{"id":48989,"date":"2026-07-11T09:48:55","date_gmt":"2026-07-11T07:48:55","guid":{"rendered":"https:\/\/www.cloudmagazin.com\/?p=48989"},"modified":"2026-07-12T22:22:29","modified_gmt":"2026-07-12T20:22:29","slug":"the-model-store-devours-the-expensive-tpu-hour","status":"publish","type":"post","link":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/","title":{"rendered":"The model store devours the expensive TPU hour"},"content":{"rendered":"<p><strong>On a TPU VM with four chips, every minute costs money. When Google made a 449-gigabyte model ready for inference after more than ten minutes in a documented benchmark, the node paid for compute power that hadn\u2019t even started calculating yet. That very first moment-when the model loads-decides the real cloud bill for AI inference on Google Kubernetes Engine.<\/strong><\/p>\n<h2>Key Takeaways<\/h2>\n<div style=\"background:#f8fbfd;border:1px solid rgba(11,183,253,0.28);border-radius:8px;padding:22px 26px;margin:16px 0 34px;box-shadow:0 1px 0 rgba(11,183,253,0.08),0 2px 8px rgba(0,0,0,0.03);\">\n<ul style=\"margin:0;padding-left:20px;line-height:1.75;color:#1a2733;\">\n<li style=\"margin-bottom:10px;\"><strong>The loading phase is the cost driver.<\/strong> For a 480-billion-parameter model, loading time on a TPU dropped from over 630 seconds to under 280 seconds once weights were streamed directly from object storage instead of taking the local detour.<\/li>\n<li style=\"margin-bottom:10px;\"><strong>The host memory requirement halves.<\/strong> The classic load path peaked at 881 gigabytes of host RAM; streaming managed with 436. The difference is budget that a node pool would otherwise have to reserve indefinitely.<\/li>\n<li><strong>The root cause lies in the architecture.<\/strong> TPU nodes have no local SSDs. The classic route temporarily consumes twice the model size in RAM. Ignoring this means paying for over-provisioning and sluggish scaling.<\/li>\n<\/ul>\n<\/div>\n<p style=\"font-size:0.88em;color:#666;margin:20px 0 32px 0;border-top:1px solid #e5e5e5;border-bottom:1px solid #e5e5e5;padding:10px 0;\"><span style=\"font-family:'SF Mono','Monaco','Consolas',monospace;color:#0879aa;font-weight:700;text-transform:uppercase;font-size:0.72em;letter-spacing:0.14em;margin-right:14px;\">Related:<\/span><a href=\"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/07\/kubernetes-finops-the-levers-that-close-70-percent-of-cluster-waste\/\" style=\"color:#333;text-decoration:underline;\">Kubernetes FinOps: The Levers Against 70 Percent Cluster Waste<\/a>&nbsp;&nbsp;<span style=\"color:#ccc;\">\/<\/span>&nbsp;&nbsp;<a href=\"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/06\/4-percent-in-the-data-center-54-at-the-power-plant-where-nvidias-water-promise\/\" style=\"color:#333;text-decoration:underline;\">4 Percent in the Data Center, 54 at the Power Plant<\/a><\/p>\n<h2 style=\"margin-top:64px;margin-bottom:20px;padding-top:16px;\">Why the Expensive TPU Waits Before It Even Computes<\/h2>\n<p>An accelerator costs per hour whether it\u2019s generating tokens or merely shoving a model into memory. With small models the overhead is invisible. With a model whose weights span hundreds of gigabytes, the loading step becomes the longest phase in a pod\u2019s life cycle.<\/p>\n<p>That directly undermines autoscaling. When a cluster scales up during peak load, every new pod must finish loading its model before it can answer the first request. If that takes several minutes, a dilemma appears: either scaling reacts too slowly and users wait, or the team keeps expensive spare capacity permanently warm. Both paths burn TPU hours.<\/p>\n<p>The math is uncomfortable. A node that loads for ten minutes and then runs for an hour sacrifices roughly one-seventh of its paid runtime to a process that produces not a single token.<\/p>\n<h2 style=\"margin-top:64px;margin-bottom:20px;padding-top:16px;\">The Hidden Memory Surcharge at Startup<\/h2>\n<p>The classic TPU load path routes through the host\u2019s RAM. The model is first read entirely into CPU memory, split there for the chips, and only then transferred. For PyTorch models without specialized load logic, this creates a memory spike of about twice the model size because the checkpoint and the prepared copy briefly coexist.<\/p>\n<p>That double-booking is the real cost center. A node pool must be sized for the spike, not normal operation. Budget is reserved for a moment that occurs only at launch.<\/p>\n<div style=\"background:#004a59;border:1px solid rgba(11,183,253,0.4);border-radius:10px;padding:34px 26px;margin:36px 0;text-align:center;box-shadow:0 0 24px rgba(11,183,253,0.08);\">\n<div style=\"font-family:'SF Mono','Monaco','Consolas',monospace;color:#0bb7fd;font-size:0.72em;font-weight:700;text-transform:uppercase;letter-spacing:0.18em;margin-bottom:12px;\">\/\/ Memory spike at load<\/div>\n<div style=\"color:#fff;font-size:2.4em;font-weight:800;line-height:1.1;\">881 GB \u2192 436 GB<\/div>\n<div style=\"color:#9fc7d6;font-size:0.95em;margin-top:10px;\">Host-RAM requirement for a 449-GB model: classic load path versus direct streaming from object storage.<\/div>\n<\/div>\n<h2 style=\"margin-top:64px;margin-bottom:20px;padding-top:16px;\">TPU nodes without local SSDs force a decision<\/h2>\n<p>Unlike many GPU instances, TPU nodes lack a fast local disk from which a model could be loaded. The weights must come from object storage or attached volumes. This intensifies the trade-off between load time, storage costs, and the risk that a node will exhaust its RAM on startup.<\/p>\n<p>The workaround Google documents for this scenario is the open-source Run:ai Model Streamer. It bypasses the local detour and streams the weights in parallel from object storage straight into the chip\u2019s memory. For the general case, the documentation cites up to six-times faster load times compared with conventional methods; the TPU case measured here delivered roughly a two-times improvement. TPU support in vLLM begins with version 0.18.0.<\/p>\n<p>Key context: the measured two-times jump applies to the described scenario of large PyTorch models. Models that already use optimized loading logic see smaller gains. If you plan to adopt the streamer, validate it against your own model type rather than copying the benchmark figure.<\/p>\n<h2 style=\"margin-top:64px;margin-bottom:20px;padding-top:16px;\">Which levers platform teams can actually pull<\/h2>\n<p>The first lever is the loading path itself. Streaming from object storage lowers the memory spike and enables smaller, cheaper node pools-two wins at once: faster pod readiness for scaling and less reserved RAM per node.<\/p>\n<p>The second lever is caching. A persistent compilation cache in object storage cuts subsequent startups noticeably because the expensive preparation step no longer runs for every pod. For zonal acceleration, a read cache can be placed in front of object storage.<\/p>\n<p>The third lever is planning. True scale-to-zero remains costly on TPUs because every cold start still pays the full load penalty. For fluctuating workloads, selective elasticity with pre-warmed reserve pods is often cheaper than pure up-and-down scaling. If you take your AI-infrastructure costs seriously-say, at the Google Cloud Summit DACH 2026-this calculation cannot be ignored.<\/p>\n<h2 style=\"padding-top:64px;margin-bottom:20px;\">Frequently Asked Questions<\/h2>\n<details>\n<summary><strong>What is the Run:ai Model Streamer?<\/strong><\/summary>\n<p style=\"margin:8px 0 4px 24px;color:#555;line-height:1.6;\">The Run:ai Model Streamer is an open-source component that loads model weights in parallel from an object store such as Cloud Storage directly into an accelerator\u2019s memory. It skips the local disk and the detour through host RAM, letting large models start faster and with a lower memory peak.<\/p>\n<\/details>\n<details>\n<summary><strong>Why is cold-start on TPU inference a cost issue?<\/strong><\/summary>\n<p style=\"margin:8px 0 4px 24px;color:#555;line-height:1.6;\">An accelerator costs per hour even while it loads a model. If loading a large model takes several minutes, scaling is delayed and teams either react slowly or keep expensive reserve capacity on standby-both of which raise the cost per successfully answered request.<\/p>\n<\/details>\n<details>\n<summary><strong>From which vLLM version does the streamer run on TPU?<\/strong><\/summary>\n<p style=\"margin:8px 0 4px 24px;color:#555;line-height:1.6;\">TPU support requires vLLM 0.18.0 or newer. For GPUs the streamer can already connect to Cloud Storage from version 0.11.1; it is enabled via an additional flag in the launch command.<\/p>\n<\/details>\n<details>\n<summary><strong>Does the two-times speed-up apply to every model?<\/strong><\/summary>\n<p style=\"margin:8px 0 4px 24px;color:#555;line-height:1.6;\">No. The measured halving of load time and memory peak refers to large PyTorch models that follow the classic path through host memory. Models with already incremental loading logic benefit less. Before switching, test with your own model.<\/p>\n<\/details>\n<details>\n<summary><strong>How does model load time relate to autoscaling?<\/strong><\/summary>\n<p style=\"margin:8px 0 4px 24px;color:#555;line-height:1.6;\">Every newly launched pod must load its model before it can answer requests. A long load phase makes auto-scaling sluggish during traffic spikes. Shorter load times let pods become ready faster, reduce the need for pre-warmed reserve capacity, and improve the utilization of paid accelerators.<\/p>\n<\/details>\n<h2>Editor\u2019s Reading Picks<\/h2>\n<ul style=\"list-style:none;padding:0;margin:16px 0 34px;\">\n<li style=\"padding:10px 0;border-bottom:1px solid #eee;\"><a href=\"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/06\/enefg-amendment-eases-data-center-efficiency-rules\/\" style=\"color:#0879aa;text-decoration:none;font-weight:600;\">Higher PUE limits and longer deadlines: how the EnEfG amendment relaxes data-centre obligations<\/a><\/li>\n<li style=\"padding:10px 0;border-bottom:1px solid #eee;\"><a href=\"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/06\/ai-integration-goes-enterprise-grade-model-context-protocol\/\" style=\"color:#0879aa;text-decoration:none;font-weight:600;\">AI integration goes enterprise-ready: the Model Context Protocol under the Linux Foundation<\/a><\/li>\n<li style=\"padding:10px 0;\">Near top-tier performance, cheaper and trained across three continents<\/li>\n<\/ul>\n<p style=\"margin:44px 0 12px;font-size:0.72em;font-weight:700;text-transform:uppercase;letter-spacing:0.18em;color:#666;\">More from the MBF Media Network<\/p>\n<div style=\"padding:14px 18px;border-left:3px solid #202528;background:#fafafa;margin-bottom:6px;\">\n<div style=\"font-size:0.7em;font-weight:700;color:#202528;text-transform:uppercase;letter-spacing:0.12em;margin-bottom:4px;\">mybusinessfuture<\/div>\n<p><a href=\"https:\/\/mybusinessfuture.com\/ki-mittelstand-pilot-skalierung-roi-roadmap\/\" style=\"font-weight:600;line-height:1.4;color:#1a1a1a;text-decoration:none;\">Scaling AI in mid-market firms from pilot to ROI roadmap<\/a><\/p>\n<\/div>\n<div style=\"padding:14px 18px;border-left:3px solid #d65663;background:#fafafa;margin-bottom:6px;\">\n<div style=\"font-size:0.7em;font-weight:700;color:#d65663;text-transform:uppercase;letter-spacing:0.12em;margin-bottom:4px;\">digital chiefs<\/div>\n<p><a href=\"https:\/\/www.digital-chiefs.de\/ki-check-kassensturz-unternehmen\/\" style=\"font-weight:600;line-height:1.4;color:#1a1a1a;text-decoration:none;\">Three AI budgets, one bottom line: no shared accounting<\/a><\/p>\n<\/div>\n<div style=\"padding:14px 18px;border-left:3px solid #69d8ed;background:#fafafa;margin-bottom:6px;\">\n<div style=\"font-size:0.7em;font-weight:700;color:#69d8ed;text-transform:uppercase;letter-spacing:0.12em;margin-bottom:4px;\">securitytoday<\/div>\n<p><a href=\"https:\/\/www.securitytoday.de\/2026\/07\/10\/was-ist-iso-27001\/\" style=\"font-weight:600;line-height:1.4;color:#1a1a1a;text-decoration:none;\">What is ISO 27001? Definition, certification and boundaries<\/a><\/p>\n<\/div>\n<p style=\"text-align:right;color:#868e96;font-size:0.85em;margin-top:48px;\"><em>Image source: AI-generated (July 2026)<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"In GKE with TPUs, model loading is the most expensive moment during AI inference. Why cold starts and host memory drive up the cloud bill.","protected":false},"author":31,"featured_media":48781,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_yoast_wpseo_meta-robots-noindex":"","_yoast_wpseo_meta-robots-nofollow":"","_yoast_wpseo_meta-robots-adv":"","_yoast_wpseo_canonical":"","_yoast_wpseo_opengraph-title":"","_yoast_wpseo_opengraph-description":"","_yoast_wpseo_opengraph-image":"","_yoast_wpseo_opengraph-image-id":0,"_yoast_wpseo_twitter-title":"","_yoast_wpseo_twitter-description":"","_yoast_wpseo_twitter-image":"","_yoast_wpseo_twitter-image-id":0,"pre_headline":"","bildquelle":"","teasertext":"","language":"de","_evm_translation_lang":"","featured_post":0,"featured_post_sortierung":0,"_wp_old_slug":[],"footnotes":""},"categories":[10],"tags":[],"industry":[],"class_list":["post-48989","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-firmensuccess"],"evm_reading_time_minutes":7,"wpml_language":"en","wpml_translation_of":48780,"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.9 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>The model store devours the expensive TPU hour<\/title>\n<meta name=\"description\" content=\"AI inference on GKE with TPUs? Model loading is the costliest step. Discover why cold starts and host memory drive up cloud bills.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"The model store devours the expensive TPU hour\" \/>\n<meta property=\"og:description\" content=\"AI inference on GKE with TPUs? Model loading is the costliest step. Discover why cold starts and host memory drive up cloud bills.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/\" \/>\n<meta property=\"og:site_name\" content=\"cloudmagazin\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/cloudmagazincom\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-11T07:48:55+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-12T20:22:29+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/07\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1792\" \/>\n\t<meta property=\"og:image:height\" content=\"1024\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Alec Chizhik\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@cloudmagazin\" \/>\n<meta name=\"twitter:site\" content=\"@cloudmagazin\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Alec Chizhik\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"5 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"NewsArticle\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/\"},\"author\":{\"name\":\"Alec Chizhik\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#\\\/schema\\\/person\\\/ce38baaa19a580268aedce096597eb3c\"},\"headline\":\"The model store devours the expensive TPU hour\",\"datePublished\":\"2026-07-11T07:48:55+00:00\",\"dateModified\":\"2026-07-12T20:22:29+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/\"},\"wordCount\":1090,\"publisher\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg\",\"articleSection\":[\"Success Stories\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/\",\"url\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/\",\"name\":\"The model store devours the expensive TPU hour\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg\",\"datePublished\":\"2026-07-11T07:48:55+00:00\",\"dateModified\":\"2026-07-12T20:22:29+00:00\",\"description\":\"AI inference on GKE with TPUs? Model loading is the costliest step. Discover why cold starts and host memory drive up cloud bills.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/#primaryimage\",\"url\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg\",\"contentUrl\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg\",\"width\":1792,\"height\":1024,\"caption\":\"Cold Start: Vom Modell-Setup zur steigenden Kostenkurve.\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/2026\\\/07\\\/11\\\/the-model-store-devours-the-expensive-tpu-hour\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/home\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"The model store devours the expensive TPU hour\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#website\",\"url\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/\",\"name\":\"cloudmagazin\",\"description\":\"Inspiration f\u00fcr Businessentscheider\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#organization\",\"name\":\"cloudmagazin\",\"url\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2020\\\/04\\\/cloudmagazin-logo-klein_menu.jpg\",\"contentUrl\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2020\\\/04\\\/cloudmagazin-logo-klein_menu.jpg\",\"width\":150,\"height\":150,\"caption\":\"cloudmagazin\"},\"image\":{\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#\\\/schema\\\/logo\\\/image\\\/\"},\"sameAs\":[\"https:\\\/\\\/www.facebook.com\\\/cloudmagazincom\\\/\",\"https:\\\/\\\/x.com\\\/cloudmagazin\",\"https:\\\/\\\/www.linkedin.com\\\/showcase\\\/cloudmagazin\\\/\"]},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/#\\\/schema\\\/person\\\/ce38baaa19a580268aedce096597eb3c\",\"name\":\"Alec Chizhik\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2026\\\/03\\\/alec-chizhik.jpg\",\"url\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2026\\\/03\\\/alec-chizhik.jpg\",\"contentUrl\":\"https:\\\/\\\/www.cloudmagazin.com\\\/wp-content\\\/uploads\\\/2026\\\/03\\\/alec-chizhik.jpg\",\"caption\":\"Alec Chizhik\"},\"description\":\"Alec is the Chief Digital Officer at Evernine and writes about cloud architectures, IT security, and digital operations practices.\",\"sameAs\":[\"https:\\\/\\\/www.linkedin.com\\\/in\\\/alecchizhik\\\/\"],\"url\":\"https:\\\/\\\/www.cloudmagazin.com\\\/en\\\/author\\\/alec\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"The model store devours the expensive TPU hour","description":"AI inference on GKE with TPUs? Model loading is the costliest step. Discover why cold starts and host memory drive up cloud bills.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/","og_locale":"en_US","og_type":"article","og_title":"The model store devours the expensive TPU hour","og_description":"AI inference on GKE with TPUs? Model loading is the costliest step. Discover why cold starts and host memory drive up cloud bills.","og_url":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/","og_site_name":"cloudmagazin","article_publisher":"https:\/\/www.facebook.com\/cloudmagazincom\/","article_published_time":"2026-07-11T07:48:55+00:00","article_modified_time":"2026-07-12T20:22:29+00:00","og_image":[{"width":1792,"height":1024,"url":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/07\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg","type":"image\/jpeg"}],"author":"Alec Chizhik","twitter_card":"summary_large_image","twitter_creator":"@cloudmagazin","twitter_site":"@cloudmagazin","twitter_misc":{"Written by":"Alec Chizhik","Est. reading time":"5 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"NewsArticle","@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/#article","isPartOf":{"@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/"},"author":{"name":"Alec Chizhik","@id":"https:\/\/www.cloudmagazin.com\/en\/#\/schema\/person\/ce38baaa19a580268aedce096597eb3c"},"headline":"The model store devours the expensive TPU hour","datePublished":"2026-07-11T07:48:55+00:00","dateModified":"2026-07-12T20:22:29+00:00","mainEntityOfPage":{"@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/"},"wordCount":1090,"publisher":{"@id":"https:\/\/www.cloudmagazin.com\/en\/#organization"},"image":{"@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/#primaryimage"},"thumbnailUrl":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/07\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg","articleSection":["Success Stories"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/","url":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/","name":"The model store devours the expensive TPU hour","isPartOf":{"@id":"https:\/\/www.cloudmagazin.com\/en\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/#primaryimage"},"image":{"@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/#primaryimage"},"thumbnailUrl":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/07\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg","datePublished":"2026-07-11T07:48:55+00:00","dateModified":"2026-07-12T20:22:29+00:00","description":"AI inference on GKE with TPUs? Model loading is the costliest step. Discover why cold starts and host memory drive up cloud bills.","breadcrumb":{"@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/#primaryimage","url":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/07\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg","contentUrl":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/07\/das-modell-laden-frisst-die-teure-tpu-stunde-cover-hero.jpg","width":1792,"height":1024,"caption":"Cold Start: Vom Modell-Setup zur steigenden Kostenkurve."},{"@type":"BreadcrumbList","@id":"https:\/\/www.cloudmagazin.com\/en\/2026\/07\/11\/the-model-store-devours-the-expensive-tpu-hour\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.cloudmagazin.com\/en\/home\/"},{"@type":"ListItem","position":2,"name":"The model store devours the expensive TPU hour"}]},{"@type":"WebSite","@id":"https:\/\/www.cloudmagazin.com\/en\/#website","url":"https:\/\/www.cloudmagazin.com\/en\/","name":"cloudmagazin","description":"Inspiration f\u00fcr Businessentscheider","publisher":{"@id":"https:\/\/www.cloudmagazin.com\/en\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.cloudmagazin.com\/en\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.cloudmagazin.com\/en\/#organization","name":"cloudmagazin","url":"https:\/\/www.cloudmagazin.com\/en\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.cloudmagazin.com\/en\/#\/schema\/logo\/image\/","url":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2020\/04\/cloudmagazin-logo-klein_menu.jpg","contentUrl":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2020\/04\/cloudmagazin-logo-klein_menu.jpg","width":150,"height":150,"caption":"cloudmagazin"},"image":{"@id":"https:\/\/www.cloudmagazin.com\/en\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/cloudmagazincom\/","https:\/\/x.com\/cloudmagazin","https:\/\/www.linkedin.com\/showcase\/cloudmagazin\/"]},{"@type":"Person","@id":"https:\/\/www.cloudmagazin.com\/en\/#\/schema\/person\/ce38baaa19a580268aedce096597eb3c","name":"Alec Chizhik","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/03\/alec-chizhik.jpg","url":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/03\/alec-chizhik.jpg","contentUrl":"https:\/\/www.cloudmagazin.com\/wp-content\/uploads\/2026\/03\/alec-chizhik.jpg","caption":"Alec Chizhik"},"description":"Alec is the Chief Digital Officer at Evernine and writes about cloud architectures, IT security, and digital operations practices.","sameAs":["https:\/\/www.linkedin.com\/in\/alecchizhik\/"],"url":"https:\/\/www.cloudmagazin.com\/en\/author\/alec\/"}]}},"_links":{"self":[{"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/posts\/48989","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/users\/31"}],"replies":[{"embeddable":true,"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/comments?post=48989"}],"version-history":[{"count":1,"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/posts\/48989\/revisions"}],"predecessor-version":[{"id":48990,"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/posts\/48989\/revisions\/48990"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/media\/48781"}],"wp:attachment":[{"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/media?parent=48989"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/categories?post=48989"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/tags?post=48989"},{"taxonomy":"industry","embeddable":true,"href":"https:\/\/www.cloudmagazin.com\/en\/wp-json\/wp\/v2\/industry?post=48989"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}