Cerebras Systems, a maker of specialized AI chips, has struck a deal to supply Gimlet Labs with its new CS-4 processors for a large-scale cloud computing service. The rollout, which will draw about 100 megawatts of power, is expected to take one to two years, with Gimlet planning to offer the systems to customers in 2027.
Gimlet Labs, an AI infrastructure startup, intends to use the chips to sell "inference" computing — the process of running AI models to generate responses, such as when a chatbot answers a question. The company's CEO, Zain Asgar, told Reuters that the target use cases include voice applications, cybersecurity, and financial services.
What is inference computing?
AI development has traditionally focused on training — the phase where models learn from vast amounts of data. But as AI products become mainstream, companies are spending more time and money on inference, which is what happens when a model is actually used. Every time a user asks a chatbot for a summary or a voice assistant to set a reminder, that's inference.
Inference is often more demanding than training in real-world settings because it needs to happen quickly and reliably, especially for applications like voice assistants that require near-instant responses. This has created a growing market for specialized hardware and cloud services designed to handle inference workloads efficiently.
Cerebras, known for its wafer-scale chips that are among the largest in the industry, has been positioning itself as a challenger to Nvidia, the dominant player in AI chips. The CS-4 is Cerebras's latest offering, and the deal with Gimlet marks a significant commercial win for the company as it seeks to expand beyond its earlier focus on training.
Why this deal matters
The agreement underscores a broader shift in the AI industry. As more businesses deploy AI tools, the demand for inference computing is rising, and cloud providers are racing to build infrastructure to meet it. Gimlet is betting that there's room for a cloud service that specializes in fast, predictable inference for AI-first companies, rather than the general-purpose clouds offered by giants like Amazon, Microsoft, and Google.
For Cerebras, the deal is a vote of confidence in its technology and its ability to compete in the inference market. The company has been working to diversify its customer base and prove that its chips can handle a variety of AI workloads, not just training.
The 100-megawatt scale is notable. To put it in perspective, a typical data center might use a few tens of megawatts, so this is a substantial deployment. It suggests Gimlet is planning for serious capacity, likely to attract large enterprise customers.
What it means for investors
For everyday investors, this deal is a reminder that the AI boom is not just about the companies building the models, but also about the infrastructure that powers them. The race to build AI clouds and specialized chips is creating opportunities and risks across the tech sector.
Companies like CoreWeave, which recently raised billions to expand its AI cloud, and Akamai, which is getting a boost from a major cloud deal, are part of this trend. The demand for AI infrastructure is also showing up in broader market moves, as chipmakers have helped drive gains in the S&P 500.
However, investors should be cautious. The AI infrastructure buildout is capital-intensive, and there's no guarantee that all the players will succeed. The market for inference computing is still young, and competition is fierce. While Cerebras and Gimlet are making bold moves, they face established rivals with deep pockets.
For those watching the sector, the key metrics to track will be how quickly Gimlet can bring its cloud online, whether it can sign up customers, and how Cerebras's chips perform in real-world inference tasks. The 2027 launch date gives both companies time to refine their offerings, but it also means investors will have to wait to see if the bet pays off.
In the meantime, the deal highlights the growing importance of inference in the AI ecosystem. As more applications move from research to production, the companies that can provide reliable, fast, and cost-effective inference will be well-positioned to benefit.

