IA 360
Current Affairs

AMD puts MI350 into production, readies Helios for 2026

AMD has begun shipping its Instinct MI350 accelerators, with up to 288 GB of memory per GPU, and previewed Helios, a rack-scale system for MI400 chips planned for 2026. The company aims to establish itself as an open alternative to Nvidia in AI data centers.

4 min read AI-generated Leer en español
AMD puts MI350 into production, readies Helios for 2026

On June 12, 2025, AMD announced MI350 production and previewed the Helios rack platform. The original source supports the documentary core of the event; vendor specifications and dates are not independent performance or identical availability for every customer.

The immediate news is the MI350X and MI355X, which AMD says are available through its server and cloud partners. AMD says the new generation delivers up to three times the AI computing capacity of the MI300 series and significantly improves inference performance—the stage at which a trained model responds to real user requests.

More memory for models that don’t fit on a single GPU

AMD’s central argument is not just computing power. Each MI350 comes with up to 288 GB of HBM3E memory, a high-speed memory technology integrated alongside the graphics processor. That figure matters because large language models, long context windows and applications serving many users need to store large amounts of data close to the GPU. Document supporting the figure.

When a model does not fit on a single GPU, it has to be split across several GPUs, which must constantly coordinate the exchange of information. That operation adds cost, power consumption and latency. More memory per accelerator can simplify deployment, especially for companies running open-source models or their own systems rather than relying solely on an external API.

The MI350 chips are based on AMD’s CDNA 4 architecture and target both training and inference. Inference has become the industry’s most important battleground: training a large model is expensive and occasional, while serving millions of queries every day creates a permanent load for data centers.

Helios takes the competition from chips to the full rack

AMD also showed off Helios, its proposal for 2026. It will not be a standalone GPU but a rack architecture bringing together Instinct MI400 accelerators, next-generation EPYC processors codenamed Venice, and Pensando networking technology.

The move reflects how major AI companies have changed the way they buy infrastructure. Customers no longer compare one chip with another; they buy entire racks containing dozens of accelerators, high-speed interconnects, storage, cooling and software. Nvidia has turned that integrated approach into an advantage with its DGX systems and Blackwell racks. Helios is AMD’s answer to that model.

The company also says the design will be built around open standards. That promise matters to data center operators that want to combine components from different vendors and avoid making their entire infrastructure dependent on a single manufacturer. But it also presents a challenge: openness must translate into systems that are easy to install, maintain and scale—not just compatibility on paper.

ROCm 7, the piece that will determine adoption

Hardware alone is not enough to displace an established platform. Nvidia maintains its dominant position in part because of CUDA, its programming ecosystem and libraries for accelerating AI workloads. AMD needs its ROCm alternative to prove reliable for researchers, developers and operations teams.

That is why the company introduced ROCm 7, the next version of its open software stack. AMD promises performance improvements and broader compatibility with leading AI frameworks and models. The practical question will be how much work it takes to move projects built for CUDA to servers with Instinct chips, and whether its debugging, monitoring and optimization tools reach the maturity demanded by large-scale deployments.

The presence of Sam Altman, OpenAI’s CEO, alongside Lisa Su at the event reflects the importance of expanding the computing capacity available to the leading AI labs.

The MI350 chips allow AMD to compete now in the current generation of AI servers. Helios and the MI400 series will determine whether it can also compete for the infrastructure contracts to be decided for 2026, when a chip’s performance will be only one part of the decision.

Turning the headline into a check

Nameplate power is not useful compute. Between a grid connection and a running model sit substations, cooling, networks, memory, storage, software and equipment availability. vendor specifications and dates are not independent performance or identical availability for every customer. An announcement becomes capacity only when every link has a date, owner and measurement.

System comparisons require a fixed workload. Training, tuning and inference use memory, communication and precision differently. A vendor maximum may depend on formats, software or models that do not match real work. To assess how to compare accelerators as complete systems rather than a row of maximums, run the same set with versions and consumption recorded.

Utilization reveals the distance between inventory and result. Installed equipment may wait for data, networking or repairs. A useful record tracks available hours, completed work, failures, total energy and delays. It also separates an operating first phase from announced final scale; mixing them turns the future into the present.

What the record must preserve

Physical cost is measured, not inferred from one magnitude. Electricity source and timing, cooling, water, construction, redundancy and service life all matter. An assessment should publish its boundary and avoid vague offsets. That inventory enables comparison and asks what capability is obtained for the resource consumed.

An evidence sheet separates four columns: what the source claims, what it shows, what it did not measure and what would change the conclusion. That discipline prevents an absence from becoming a promise and a condition from vanishing in summary. It also lets the story be updated without rewriting history from a later outcome.

Include a negative case before deciding. Find a situation where the system, rule, transaction or study does not meet the need and record the signal that would require stopping. Selected successes show that something can happen; the negative case reveals the boundary and lowers the cost of discovering it after deployment.

The skill that outlasts the announcement

A valid comparison preserves denominator and axis. It does not pit a point figure against an average, future capacity against installed capacity or a forecast against an observation. When two sources use similar language, reconstruct what they counted and over what period. If those differ, publish them as different measures instead of inventing a ranking.

The record should survive a version change. Keep URL, consultation date, document, configuration and decision. When new evidence appears, add it with its date and explain what it changes. That traceability prevents opposite errors: keeping an expired conclusion or pretending later information was known on the event date.

The transferable skill in this story is how to compare accelerators as complete systems rather than a row of maximums. The procedure is short: name the document, preserve the date, fix the axis, find the condition and design a check that can fail. With those steps, a reader need not accept or reject the announcement by intuition; the decision follows a visible chain of evidence.

Before closing, another person should be able to reconstruct the conclusion without knowing the headline. Give them the sources, conditions and negative case, then ask what they would accept and reject. If they need an assumed intent, a figure without a denominator or an undated later fact, the chain still has a gap. That short review catches errors that fluent prose can conceal.

The result is not a permanent score but a dated, revisable decision. Set when to measure again and which signal triggers an earlier review. Caution then does not paralyze; it turns uncertainty into a monitoring condition. It also prevents an announcement from receiving credit for a later improvement that was not available when the decision was made.

Finally, preserve the alternative. The question is not only whether the announcement works, but whether it improves the process compared with keeping the current approach, using another tool or waiting for evidence. A concrete baseline prevents novelty from being mistaken for benefit. The decision may be to proceed, limit scope or make no change yet, always for a checkable reason.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close