← Back

Cerebras Runs Google's Gemma 4 at 1,500 Tokens/Sec to Melt Your Screens

Original version ·

While big tech keeps trying to convince us that a three-second delay for a simple chatbot response is 'magic', some mad lads just strapped rocket boosters to open-weights AI and gave it eyes. This is what actual hardware wizardry looks like when you stop playing nice.

The silicon-growing monster Cerebras, famous for building microchips the size of an entire dinner plate, has unleashed its hardware on Google's latest creation. During a private run, their system processed the Gemma 4 31B model so fast that the text doesn't just scroll—it basically teleports onto the screen. To make things even more chaotic, they also unlocked multimodal capabilities, meaning this speed demon can now look at images and screenshots instead of just reading boring text files.

Usually, these specialized wafer-scale engines behave like sports cars that only work on perfectly flat roads, preferring sparse, specialized calculations. However, Google's model is a heavy, dense beast without any of those fancy Mixture-of-Experts shortcuts. Forcing this giant chunk of silicon to process a dense model at warp speed is the hardware equivalent of making a freight train drift around a tight corner.

The creators are already flexing the practical use cases, focusing on tasks where waiting even five seconds feels like an eternity. They demonstrated the system digesting bloated corporate dashboards instantly, summarizing massive PDF reports before a human can even take a sip of coffee, and performing instant debugging where a developer feeds it a broken UI screenshot and receives a ready-to-paste code patch. The goal is clearly to convince developers to abandon their current cloud providers and migrate their applications immediately.

When stacked against the industry standard Claude Haiku from Anthropic, the plate-sized chip delivers a fifteen-fold speed increase. On benchmark tests, the model scores nearly identically to its rival, meaning users get the exact same quality of output but without the awkward, elevator-music pause while the server farm thinks. Public access is scheduled to open up by the end of the month.

AI companies have spent years conditioning everyone to accept lag as a natural law of computing. Watching a giant slab of silicon obliterate that bottleneck proves that the real limitation isn't the intelligence of the models, but the corporate refusal to build better engines.

Source: Cerebras

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

0/24
  1. No comments yet.