TL;DR · 30-second read
The Short Version
Anthropic, the company behind the Claude chatbot, is starting to design its own computer chips for running its AI. Samsung is reportedly in talks to build them.
Today most AI runs on chips from Nvidia, which are expensive and use a lot of electricity. Google, Amazon, Meta, Microsoft and OpenAI already make their own chips.
What most people miss is that these custom chips are built to answer people’s questions, not to create new AI models. Creating new models still mostly happens on Nvidia hardware. Nvidia’s rivals are winning ground on the everyday work of answering requests.
Anthropic is building an in-house chip development team to co-design custom ASICs (application-specific integrated circuits, chips built for one kind of job) to run its Claude models. The company told Business Insider it is co-designing the chips so that Claude can run faster and more efficiently at the scale its customers need. It is hiring engineers now, and the job listing stresses delivering the design on schedule. Tom’s Hardware reported the development on August 7, 2026. It added that The Information had reported in July that Anthropic was in talks with Samsung to manufacture the chips.
Anthropic has not named its design partner. No chip specifications, production volumes or deployment dates have been disclosed.
Executive Summary
Anthropic is the latest frontier AI developer to pursue its own silicon. Google, Amazon, Meta, Microsoft and OpenAI already have custom chip programs. The chips are aimed at inference, which is the work of running a trained model to answer requests. They are not aimed at training, the far more compute-intensive process of building the model in the first place.
That distinction matters more than the headline rivalry with Nvidia. Training advanced models still depends heavily on Nvidia GPUs. Inference can run on a much wider range of hardware, and its cost recurs with every prompt a customer sends. For a company whose usage is growing quickly, especially from agentic AI that consumes many more tokens per task than a single chat exchange, inference is the line item that scales. It is therefore the one worth engineering down.
The reported Samsung link adds a supply-chain dimension. Samsung is a small player in custom chip co-design, but it brings manufacturing capacity and memory. Memory is one of the scarcest inputs in AI hardware right now.
Why Nvidia’s Exposure Sits in Inference, Not Training
Every custom AI chip program disclosed so far aims first at running models, not building them. The economics explain why. Training a frontier model is an episodic, enormous job, and Nvidia’s GPUs and software ecosystem remain the default for it. Inference is continuous. Each question a user asks, and each step an AI agent takes, consumes compute. As usage grows, inference spending grows with it. Agentic workloads multiply the token count per task, so the growth is faster still.
Inference also tolerates specialization. A chip that only needs to execute a known family of models efficiently can drop much of a general-purpose GPU’s flexibility. In exchange it gains lower power draw and lower cost per request. Estimates put custom silicon’s total cost of ownership advantage at up to 65% compared with Nvidia GPUs. That figure is a ceiling across the industry, not a disclosed number for Anthropic’s chip, which has no published specifications. Anthropic’s own stated goal is narrower and consistent with this reasoning: making Claude run faster and more efficiently at customer scale.
For Nvidia, this means the pressure is concentrated in inference deployments. Each custom chip installed for inference displaces a GPU that would otherwise have done that work. Training, where Nvidia’s position is strongest, is not what these programs target. The large AI labs are also unlikely to become fully self-sufficient. Demand is growing too fast, and Nvidia retains strong control over access to the supply chain. The realistic outcome is mixed fleets, with GPUs for training and flexible workloads and in-house chips for high-volume serving.
The Real Cost Is the Software, and the Years
Designing a chip is only the visible part of the effort. A custom accelerator needs a full software stack, including compilers, runtimes and kernels, before it can run production models. The models themselves need tuning to exploit the hardware. That burden is why most companies sensibly buy general-purpose GPUs. Custom silicon only pays back at very large scale, where small per-request savings multiply across billions of requests.
The reference point is sobering. Google has built its Tensor Processing Units for 12 years, with Broadcom as co-design partner on each generation. Amazon runs parallel Trainium and Inferentia lines. Meta has announced several MTIA designs for deployment through 2027. Anthropic is at the hiring stage, and its first chip will arrive into a market where rivals are several generations in. The job listing’s emphasis on schedule suggests the company understands this, but no delivery target has been disclosed.
Samsung, Memory and a Concentrated Co-Design Market
Custom AI chips are rarely designed alone. Broadcom and Marvell account for roughly 95% of the ASIC co-design market. Broadcom works with Google, OpenAI and Meta, as well as ByteDance and Fujitsu. It claims a $73 billion order backlog and expects more than $100 billion in annual AI chip revenue by the end of 2027. Marvell supports Amazon’s Trainium and Microsoft’s Maia and is expected to earn upward of $11 billion from co-design work in 2026. Whichever partner Anthropic has chosen for design, that partner is serving a crowded order book.
Samsung’s reported role is manufacturing, and its appeal is not mainly design share. Samsung would bring fabrication capacity and chip design expertise. It would also bring access to memory, which is in short global supply. Modern AI accelerators depend on high-bandwidth memory (HBM) stacked alongside the processor. That memory, along with advanced packaging such as TSMC’s CoWoS, has been a practical limit on how many AI chips the industry can ship. A manufacturing partner that also makes memory could ease one of the tightest constraints a new chip program faces. Neither company has confirmed such an arrangement.
What Changes for Data Center Operators
For the facilities that host this hardware, custom inference chips add variety to planning. GPUs are power-hungry, and inference ASICs are built specifically to cut power per request. A fleet that mixes Nvidia GPUs with several in-house accelerators could bring different rack densities, cooling needs and networking layouts to the same campus. Operators serving the large AI labs should expect less standardization over the next few hardware cycles, not more.
Background
Anthropic develops the Claude family of AI models and sells access to consumers, businesses and government customers. It has grown quickly over the past year, and agentic AI, where models carry out multi-step tasks, has sharply increased the volume of computation each customer uses. Like other frontier labs, Anthropic runs its models on hardware from several suppliers, with Nvidia GPUs the industry’s default.
Custom AI silicon has become standard practice among the largest AI companies and cloud providers, often called hyperscalers because of the scale of their data center fleets. The market for designing these chips is dominated by Broadcom and Marvell. Much of the manufacturing and advanced packaging runs through TSMC, and memory supply has become a key constraint. Source: Anthropic co-designing custom AI inference chips to bypass costly Nvidia GPUs — Samsung reported as manufacturing partner for Claude maker (Tom’s Hardware): Tom’s Hardware’s report on Anthropic’s in-house inference chip team and its reported manufacturing talks with Samsung.Sources

