AI inference platform Baseten is nearing a funding round of roughly $1.5 billion, according to a June 19, 2026 report from PYMNTS. The report ties the raise directly to surging demand for inference — the work of running trained AI models in production — rather than for model training.
Terms, investors, and valuation were not detailed in the headline-level report, and the round had not been confirmed as closed at publication time.
Executive Summary
According to the report, Baseten — a company that helps businesses deploy and serve AI models at scale — is close to raising approximately $1.5 billion in new capital. For a company that was a mid-sized startup only two years earlier, a raise of this magnitude would rank among the largest ever for a dedicated inference provider.
The significance is less about one company than about where AI infrastructure money is now flowing. For the first few years of the generative-AI boom, capital chased training: the enormous one-time compute jobs that create frontier models. A $1.5 billion round for an inference specialist signals that investors now see the recurring, usage-driven business of serving models to end users as the larger and more durable prize.
That said, the source is thin. A single report of a round that is ‘near’ closing establishes investor intent and market temperature, but not final terms, valuation, or how the money will be spent. Those distinctions matter for anyone reading this as a market signal.
Inference Becomes the Center of Gravity
Training a large AI model is a one-time capital event; inference is a bill that arrives every time anyone uses the model. As AI applications have moved from demos into daily production use, the aggregate compute spent answering queries has grown continuously, while training runs remain episodic and concentrated among a handful of frontier labs. A near-$1.5 billion bet on an inference specialist is a bet that this recurring workload — not the headline-grabbing training runs — is where sustained revenue accumulates.
This inversion matters for the whole infrastructure stack. Training clusters favor a few gigantic, tightly coupled GPU installations. Inference favors distributed capacity closer to users, high utilization, and relentless cost-per-token optimization. If the money is following inference, demand patterns for data center capacity, networking, and power will follow it too.
Why Inference Platforms Command This Kind of Capital
Inference sounds simple — run the model, return the answer — but doing it profitably at scale is an engineering discipline of its own: batching requests, compiling models to specific chips, autoscaling against spiky traffic, and squeezing latency low enough for real-time products. Companies like Baseten sell that discipline as a service, sitting between raw GPU suppliers and application builders who don’t want to run their own model-serving operation.
The catch is that the business is capital-hungry in both directions. Serving customers requires reserving expensive GPU capacity ahead of demand, and competing on price requires continuous optimization investment. A $1.5 billion war chest, if the round closes as reported, is plausibly less about runway than about locking up compute supply and engineering talent before rivals do.
Winners, Losers, and the Squeeze in the Middle
The clearest beneficiaries of an inference-led cycle are the layers underneath: GPU vendors, specialized AI clouds, and the data center and power providers that host distributed serving capacity. The most exposed parties are undifferentiated middlemen — inference is a market where hyperscalers (Amazon, Google, Microsoft), well-funded independents, and open-source serving stacks all compete, and per-token prices have fallen steadily across the industry.
That competitive pressure cuts both ways for Baseten. A massive raise validates the category but also raises the stakes: the company would need to convert capital into durable advantages — proprietary optimizations, enterprise trust, sticky deployments — faster than falling inference prices erode margins. Investors appear to be betting that scale itself becomes the moat. That thesis is credible but unproven, and the report offers no revenue or margin data to test it against.
Background
Baseten was founded in 2019 in San Francisco, initially building tools that let software teams deploy machine-learning models without specialized infrastructure staff. The generative-AI boom transformed that niche into one of the industry’s fastest-growing markets, and the company raised successive venture rounds through 2025 that reportedly pushed its valuation past $2 billion.
The broader market context is a widely discussed shift in AI economics: as chatbots, coding assistants, and AI-powered products moved into everyday production use, industry attention moved from training models to serving them. Inference specialists — alongside GPU clouds and the data center operators beneath them — became prime beneficiaries of that shift, setting the stage for the mega-round reported here.
Accenture, one of the world’s largest technology consultancies, has made an investment in Dragos, a specialist in operational technology (OT) cybersecurity — the discipline of protecting the industrial control systems that run power grids, pipelines, manufacturing plants, and other critical infrastructure. Industry publication Industrial Cyber reported the move on June 19, 2026, framing it as the start of a new phase for OT security in critical infrastructure.
Financial terms and deal structure were not detailed in the source available to us, but the strategic signal is clear: a consulting giant with reach into most of the world’s largest enterprises is putting capital behind a pure-play industrial cybersecurity vendor.
Executive Summary
The announcement pairs two very different kinds of companies. Accenture sells transformation programs, managed services, and security consulting to boards and CIOs at global scale. Dragos builds software and threat intelligence focused narrowly on industrial control systems (ICS) — the programmable controllers, sensors, and safety systems that keep physical infrastructure running. An investment tie-up suggests Accenture wants OT security woven into its mainstream security offerings, and that Dragos wants distribution far beyond what a specialist sales force can reach.
Why it matters: OT security has long been treated as a niche — technically distinct from IT security, bought by plant engineers rather than CISOs, and chronically underfunded. A stamp of approval from a firm of Accenture’s size is the kind of signal that moves a category from specialist concern to standard line item in enterprise security budgets. For operators of critical infrastructure, including data centers whose power, cooling, and building-management systems are themselves OT, that shift is overdue.
The caveat: on the information available, this is a directional signal, not a quantified commitment. The size of the investment, its terms, and any joint go-to-market obligations were not disclosed in the source we reviewed, so the scale of the bet remains an open question.
Why OT Security Is Finally Going Mainstream
For decades, industrial control systems were protected mainly by isolation — the so-called air gap between plant networks and the internet. That era is over. Remote monitoring, predictive maintenance, cloud analytics, and now AI have wired factory floors and substations into corporate networks, a trend known as IT-OT convergence. Every new connection is a potential path for attackers, and ransomware crews have learned that halting physical operations creates far more pressure to pay than encrypting office files ever did.
Regulators have noticed too. Critical-infrastructure operators in the US, EU, and elsewhere face expanding incident-reporting and resilience obligations, which push OT security out of the plant manager’s discretionary budget and into board-level compliance spending. When a category becomes a compliance requirement, mainstream buyers need mainstream suppliers — which is precisely the gap a consultancy-backed specialist can fill.
The Consultancy-Plus-Specialist Playbook
The logic of the deal runs both ways. Accenture gets credible depth in a domain where generalist security practices are often thin: defending 20-year-old programmable logic controllers requires different tools, different threat intelligence, and a different tolerance for downtime than patching laptops. Dragos gets what every specialist vendor struggles to build — access to thousands of enterprise relationships and the army of delivery consultants needed to deploy and operate OT monitoring at scale.
There is also a market-structure story here. Large integrators and consultancies have been steadily aligning with, investing in, or acquiring security specialists, because customers increasingly want outcomes (‘secure my plant’) rather than products. If that pattern holds, competing OT vendors will face pressure to find their own scale partners, and independent specialists without one may find enterprise deals harder to win. The counterweight: deep consultancy alignment can make a vendor feel less neutral to customers who work with rival integrators.
What It Means for Infrastructure Operators — Including Data Centers
The ‘critical infrastructure’ framing usually evokes power utilities and pipelines, but the lesson lands closer to home for anyone running physical infrastructure. A modern data center is an OT environment: building management systems, power distribution units, generators, chillers, and fire suppression all run on industrial protocols with the same legacy-security problems as a factory floor. An attacker who compromises cooling controls can take down a facility as surely as one who breaches the servers inside it.
Mainstreaming OT security should, over time, mean more mature tooling, more available expertise, and more benchmark data for these environments. In the near term, operators should expect the opposite of relief: more auditor questions, more customer security questionnaires that now include OT sections, and more pressure to show visibility into control networks that were historically unmonitored. Getting an asset inventory of your OT environment before someone else asks for it remains the practical first step.
Background
Dragos was founded in 2016 by Robert M. Lee and colleagues with backgrounds in US government cyber operations, and built its business entirely around industrial control system defense — a deliberate contrast with generalist security vendors. It became one of the category’s flagship names, known for its OT monitoring platform, its threat-intelligence tracking of adversary groups that target industrial systems, and incident-response work on high-profile infrastructure attacks. The company reached unicorn status (a valuation above $1 billion) in 2021 as investor interest in industrial security accelerated.
Accenture is a global professional-services firm with one of the largest security consulting and managed-services practices in the world, serving most major industrial, energy, and utility companies. Its investments and acquisitions have repeatedly signaled which security categories it expects clients to spend on next — which is why a bet on OT security draws attention beyond the deal’s undisclosed size.
Federal energy regulators have approved a plan to accelerate grid interconnection for AI-focused data centers, according to reporting from The Hill dated June 18, 2026. The action is aimed at shortening the multi-year waits large new electric loads currently face before they can plug into the U.S. transmission system.
Executive Summary
The Federal Energy Regulatory Commission (FERC) — the U.S. agency that oversees interstate electricity transmission — has cleared a policy pathway to speed how quickly new AI data centers can connect to the grid. Interconnection, the technical and legal process of joining a large customer or generator to the transmission network, has become one of the tightest bottlenecks in the buildout of AI infrastructure.
The decision matters because power, not chips or real estate, is now the binding constraint on where and when hyperscale AI campuses can come online. Faster interconnection could unlock stalled projects and shift competitive dynamics among regions, utilities, and cloud providers. It also raises pointed questions about cost allocation, reliability, and fairness to existing ratepayers that the underlying reporting does not fully resolve.
Why Interconnection Became the AI Bottleneck
Modern AI training campuses can draw hundreds of megawatts — the equivalent of a small city — from a single site. Under standard interconnection procedures, utilities and regional grid operators must study how such loads affect voltage, congestion, and reliability before allowing them to energize. Those studies, layered on top of transmission upgrades that can take years to build, have produced queues stretching well beyond the planning horizon of any AI product cycle. A FERC-blessed fast-track pathway signals that regulators now view the status quo as economically untenable for a strategically important sector.
For laypeople, the shorthand is this: getting a large factory or data center plugged into the high-voltage grid is not like flipping a switch. It requires engineering studies, contracts, and sometimes new wires or substations. Cutting that timeline is powerful — and, if done badly, risky.
Winners, Losers, and Regional Reshuffling
Hyperscalers and colocation developers with shovel-ready sites near existing transmission capacity are the most obvious beneficiaries. So are utilities in regions with headroom on their networks, which can now court AI load with a credible speed-to-power pitch. Conversely, developers whose projects depended on being ahead in a strict first-come, first-served queue may see their positional advantage erode if fast-track criteria reward readiness or strategic importance over queue date.
Regional grid operators — PJM in the Mid-Atlantic, ERCOT in Texas, MISO in the Midwest, and others — will translate the federal signal into local tariffs and procedures. Expect divergence: some markets will move aggressively, others cautiously, producing a patchwork that data center site selectors will have to navigate carefully.
Reliability, Ratepayers, and the Fairness Question
Speed has trade-offs. Interconnection studies exist to protect the grid from destabilizing new loads and to fairly allocate the cost of network upgrades. Compressing that process invites two legitimate concerns: whether reliability margins are being quietly thinned, and who ultimately pays for the transmission investments that AI campuses require. If costs are socialized to residential and small-business ratepayers, expect political blowback from consumer advocates and state regulators, some of whom have already pushed back on hyperscaler-driven rate designs.
A fair reading of the policy shift is that it is neither a giveaway nor a threat on its face — the details of eligibility, cost allocation, and reliability safeguards will determine whether it holds up. Those details are precisely what the initial reporting leaves thin, and they warrant close scrutiny from all sides, including industry proponents.
Background
The U.S. electric grid was largely built for a world of predictable, gradually growing demand. The arrival of AI training and inference at scale has upended that assumption, with individual campuses requesting more power than some entire industrial parks. At the same time, transmission construction has slowed under permitting, siting, and supply-chain pressures, producing interconnection queues that in some regions exceed the total installed capacity of the grid itself.
FERC has spent recent years working through a series of reforms to modernize interconnection procedures, including changes to generator queue processing. Extending similar urgency to large loads such as AI data centers marks a notable expansion of that agenda and reflects the growing recognition that power access is now central to U.S. competitiveness in artificial intelligence.
Accenture announced on June 18, 2026 that it will strengthen critical-infrastructure defense with an end-to-end cybersecurity platform, positioning the offering as a response to AI-driven cyber threats and rising geopolitical risk. The announcement frames the platform as spanning the full defensive lifecycle for operators of essential services rather than addressing a single security niche.
The release, distributed under Accenture’s own name, provides the strategic framing — critical infrastructure, AI-era threats, geopolitics — but the public summary offers few technical or commercial specifics, so the scope of what has actually launched versus what is planned remains to be detailed.
Executive Summary
Accenture, one of the world’s largest technology consulting and managed-security providers, is moving to package its critical-infrastructure security work as a platform — a productized, presumably repeatable offering — rather than purely as bespoke consulting engagements. The stated rationale is twofold: attackers are increasingly using artificial intelligence to scale and sharpen intrusions, and geopolitical tension has made power grids, pipelines, transport networks, and communications systems more attractive targets for state-aligned actors.
Why it matters: critical infrastructure sits at the intersection of two historically separate security worlds — information technology (IT, the business systems) and operational technology (OT, the industrial control systems that physically run plants and grids). Most operators struggle to defend both coherently. An ‘end-to-end’ platform from a firm with Accenture’s reach signals that the biggest services players believe this convergence is now a mainstream market, not a specialist niche.
That said, the announcement as publicly summarized is strategic positioning more than a spec sheet. Pricing, availability, named technology components, and customer commitments are not detailed in the source material, so buyers should treat this as a statement of direction until Accenture publishes the specifics.
From Billable Hours to Platforms: A Structural Shift in Security Services
Consulting firms have traditionally sold cybersecurity as labor — assessments, incident response, staff augmentation — billed by the engagement. A ‘platform’ announcement signals a different ambition: recurring revenue, standardized tooling, and outcomes that scale beyond the headcount deployed. For Accenture, which has spent years acquiring security firms and building managed-services capacity, packaging that portfolio as an end-to-end platform is a logical next step and mirrors a broader industry pattern of services firms productizing what they previously customized.
The open question is what ‘platform’ means in practice here. The term can describe genuinely integrated software, a curated bundle of partner technologies operated by Accenture, or a branded methodology wrapping existing services. Each is legitimate, but they carry very different implications for switching costs, integration effort, and vendor lock-in. The public announcement does not yet make that distinction, and buyers should press for it.
Why Critical Infrastructure Is the Battleground of the AI Threat Era
Critical infrastructure — energy, water, transport, healthcare, communications, and the data centers underpinning all of them — is uniquely exposed because its operational technology was often built decades ago, before modern security assumptions, and cannot simply be patched or rebooted like an office laptop. Connecting those systems to modern networks created efficiency, but also a pathway for attackers. Accenture’s framing around AI-driven threats reflects a real dynamic: AI tools lower the cost of reconnaissance, phishing, and vulnerability discovery, letting attackers probe many targets at machine speed. Defenders, in turn, are looking to AI to triage alerts and spot anomalies faster than human analysts can.
The geopolitical framing is equally grounded. Governments in the US, EU, and elsewhere have spent recent years warning that state-aligned actors pre-position inside infrastructure networks, and regulation — from the EU’s NIS2 directive to US incident-reporting rules for critical sectors — is pushing operators toward demonstrable, auditable security programs. That regulatory pull, as much as the threat itself, is what creates a commercial market for end-to-end offerings.
Winners, Losers, and the Competitive Field
If Accenture executes, the pressure lands first on mid-sized OT-security specialists and regional integrators, who compete on depth but cannot match a global firm’s delivery footprint or board-level relationships. Pure-play OT security vendors may see it differently: a consultancy platform typically needs underlying detection technology, so the announcement could expand partnership channels as easily as it threatens them. Rival integrators and the security arms of large IT firms will read this as confirmation that critical-infrastructure security is consolidating into large, multi-year programs rather than point purchases.
For infrastructure operators and data-center providers, the practical takeaway is that the market is maturing toward accountability: buyers increasingly want one throat to choke across IT and OT, and large providers are positioning to be that throat. Whether a single end-to-end provider is desirable — versus a best-of-breed mix — remains a genuine architectural debate, and the right answer depends on an operator’s in-house capability, regulatory exposure, and tolerance for vendor concentration risk.
Background
Accenture is a Dublin-headquartered global professional-services firm and one of the largest cybersecurity services providers in the world, having assembled its security practice through sustained investment and a long series of acquisitions spanning incident response, managed detection, and industrial-control-system security. Its clients include large enterprises and government bodies across the sectors commonly designated as critical infrastructure.
The market context is a decade-long convergence of IT and OT security, accelerated recently by two forces: the arrival of generative AI as both an attack amplifier and a defensive tool, and heightened geopolitical tension that has put state-aligned intrusions into infrastructure networks on government agendas in the US, Europe, and Asia. Regulators have responded with binding security and incident-reporting requirements, turning what was once discretionary spending into compliance-driven demand — the commercial backdrop against which Accenture’s platform announcement lands.
Data Center Knowledge published a report on June 18, 2026 examining why the data center industry continues to rely on evaporative cooling — a heat-rejection method that consumes large volumes of water — even as public and regulatory backlash over water use intensifies. The piece frames the industry’s position as hesitation rather than refusal: operators broadly acknowledge the water problem but have been slow to abandon a technology that remains cheaper and more energy-efficient than the alternatives.
Executive Summary
The report’s core subject is a tension the industry has lived with for years and that the AI build-out has sharpened: evaporative cooling rejects heat by evaporating water, which makes it highly energy-efficient but water-hungry, while the main alternatives — dry (air-cooled) systems and refrigerant-based chillers — save water at the cost of higher electricity consumption, larger equipment footprints, or both. In markets where power is the scarcest commodity a data center can buy, trading water savings for a bigger electrical load is not a simple upgrade; it is a genuine engineering and economic trade-off.
That trade-off is why the headline speaks of hesitation. Operators face mounting pressure from drought-affected communities, local governments, and sustainability commitments to cut water use, and technologies such as closed-loop liquid cooling and hybrid systems are maturing. But retrofitting existing facilities is expensive, and for new builds the calculus depends heavily on local climate, water price, and power availability — variables that differ from one metro to the next. The result is an industry moving unevenly rather than uniformly, which is precisely the dynamic worth understanding for anyone siting capacity or evaluating operators’ sustainability claims.
The Water-for-Energy Trade at the Heart of Cooling
Every data center must move heat from chips to the outside world, and the physics offers no free option. Evaporative systems — cooling towers and their variants — exploit the fact that evaporating water absorbs enormous amounts of heat, which lets a facility reject heat with comparatively little electricity. Dry coolers and air-cooled chillers avoid consuming water but must push heat into the air mechanically, which takes more fan and compressor power, especially on hot days when the temperature difference working in the operator’s favor shrinks. In plain terms: saving water usually means burning more electricity, and in an era when grid connections are the binding constraint on data center growth, extra megawatts spent on cooling are megawatts not available for revenue-generating compute.
This is the economic logic the Data Center Knowledge piece points at with its framing of industry hesitation. An operator that switches a large campus from evaporative to dry cooling is not just paying for new equipment; it is accepting a permanently higher power draw — degrading power usage effectiveness, the industry’s standard efficiency metric — and potentially reducing the sellable IT capacity of a power-constrained site. Where water is cheap and power is scarce, the incumbent technology keeps winning on spreadsheets even as it loses in public opinion.
Why the Backlash Is Getting Harder to Price at Zero
For most of the industry’s history, water was effectively an afterthought in site selection — abundant, inexpensive, and invisible to the public. That has changed. Data center water consumption has become a recurring flashpoint in drought-prone regions, a subject of local permitting fights, and a standard line of questioning for journalists and community groups evaluating new projects. Operators now routinely publish water usage effectiveness figures and, in some cases, commit to becoming “water positive” — replenishing more water than they consume.
The practical consequence is that water carries a growing shadow price beyond the utility bill: longer permitting timelines, conditions attached to approvals, reputational exposure, and in the worst case the loss of a site altogether. The report’s premise — that the industry hesitates rather than transitions — suggests that many operators still judge those risks manageable relative to the hard costs of switching. Whether that judgment holds depends largely on how regulators and communities act next, which varies enormously by jurisdiction.
The Alternatives Are Real, but Not Drop-In
The transition options are well understood in engineering terms. Dry cooling eliminates onsite water evaporation at the cost of energy and space. Hybrid systems run dry most of the year and evaporate water only during peak heat, cutting consumption substantially without the full energy penalty. Direct-to-chip liquid cooling and immersion cooling — increasingly common in AI deployments because high-density chips demand them — move heat in closed loops that consume little or no water onsite, though the heat still has to be rejected somewhere, and that final stage can itself be wet or dry. None of these is a simple swap for an operating facility: cooling infrastructure is capital-intensive, deeply integrated with a building’s design, and typically replaced on decade-plus cycles.
That replacement cycle is the quiet variable in the whole debate. The realistic path for the industry is less about retrofitting the installed base and more about what gets designed into the enormous wave of new construction now underway. If new AI-era facilities standardize on low-water designs where climate and economics allow, the fleet’s water profile shifts over years, not quarters. If they default to evaporative cooling because power constraints dominate, the backlash the report describes is likely to intensify.
Winners, Losers, and the Siting Chessboard
The cooling transition redistributes advantage. Cooler, water-rich regions gain appeal because they make both wet and dry cooling cheaper; hot, arid markets that boomed on cheap land and power face the sharpest version of the water-versus-energy dilemma. Vendors of hybrid and liquid cooling systems benefit from every tightening of water rules. Utilities and municipalities gain leverage, since water service is becoming a negotiated element of large deals rather than a formality. And operators that invested early in low-water designs acquire a permitting and public-relations asset that is difficult for laggards to replicate quickly. Buyers of colocation and cloud capacity should read cooling architecture as a proxy for siting risk: a facility’s water dependence is now part of its long-term cost and continuity profile.
Background
Cooling is one of the two great resource demands of data centers, alongside electricity: every watt a server consumes becomes heat that must be removed. For decades, evaporative cooling towers have been a workhorse of large-scale heat rejection across many industries because evaporating water is thermodynamically cheap. Data centers adopted the approach widely as the industry scaled through the cloud era, and it helped drive the sector’s headline efficiency gains. The AI construction boom that accelerated through the mid-2020s raised the stakes on both sides of the equation — far denser computing produces far more heat, while the communities hosting these facilities have grown increasingly vocal about local water and power impacts. Trade publication Data Center Knowledge, which published the report discussed here, has tracked this cooling debate as one of the defining infrastructure questions of the AI build-out.
AI inference provider Baseten is reportedly raising $1.5 billion in new funding, according to a June 18, 2026 report from SiliconANGLE. The report describes a round in progress rather than a closed deal, and terms such as valuation, investors, and structure were not disclosed in the source material.
If the figure holds, it would rank among the largest financings yet for a company focused specifically on inference — the business of serving AI models to end users — rather than on training them.
Executive Summary
The headline fact is simple: Baseten, a platform that helps companies deploy and run AI models in production, is reported to be raising $1.5 billion. Because this is a media report of an in-progress raise rather than a company announcement, the number should be treated as provisional until confirmed.
The significance is less about one company and more about what the capital is chasing. For the past several years, the biggest checks in AI infrastructure went to training — the enormous, one-time computation of building frontier models. A ten-figure round for an inference specialist suggests investors now believe the durable, recurring revenue sits in serving models at scale, every second of every day, to real applications.
For infrastructure operators, that shift matters. Inference workloads have different economics than training: they run continuously, they are latency-sensitive, they favor geographic distribution over single giant campuses, and they reward efficiency per query rather than raw peak compute. Where the money goes, data center design, power planning, and network architecture tend to follow.
From Training to Serving: Why the Money Is Moving
Training a large AI model is a capital event — vast, concentrated, and episodic. Inference is an operating expense that scales with usage: every chatbot reply, code completion, and document summary is an inference call. As AI products mature from demos into deployed software with paying users, the volume of inference grows with adoption, and it never stops. Investors underwriting a reported $1.5 billion round are, in effect, betting that this recurring workload — not the next training run — is where sustainable revenue accumulates.
That thesis has a sound structural basis. A model is trained once but served millions or billions of times, so over a product’s life the cumulative compute spent on inference can dwarf what was spent creating the model. Companies that sit in the serving path — optimizing latency, managing GPU fleets, autoscaling with demand — collect a toll on every one of those calls.
What a War Chest Buys in the Inference Business
Inference platforms are capacity businesses as much as software businesses. To guarantee customers low latency and high availability, a provider must secure GPUs — either owned, leased from cloud providers, or contracted from specialized GPU clouds — ahead of demand. That is capital-intensive, and it is the most plausible use for a raise of this size: locking up compute supply, expanding into more regions to cut round-trip latency, and funding the engineering that squeezes more throughput out of each accelerator.
Scale also buys negotiating power. Larger committed volumes typically mean better pricing on hardware and colocation, which flows through to more competitive per-token pricing for customers. In a market where inference is increasingly bought like a commodity — priced per million tokens — cost structure is strategy.
A Crowded Field, and the Hyperscaler Question
Baseten does not operate in a vacuum. Dedicated inference providers compete with one another, with GPU-cloud operators moving up the stack, and — most importantly — with the hyperscale clouds, which bundle inference into broader platforms, and with model developers offering their own hosted APIs. The bear case for any independent inference company is that serving becomes a thin-margin utility captured by whoever owns the most silicon.
The bull case is specialization: enterprises running open-weight or fine-tuned models often want performance tuning, deployment control, and price transparency that general-purpose clouds don’t prioritize. A raise of the reported magnitude suggests at least some sophisticated investors find the bull case credible — though it is worth remembering that a reported raise reflects investor conviction, not proven unit economics. The release-level information here does not tell us Baseten’s revenue, margins, or utilization, and those are the numbers that will ultimately decide the argument.
Background
Baseten emerged in the wave of machine-learning infrastructure startups that formed as companies moved AI models out of research labs and into production applications. Its focus is the deployment layer: rather than training models or selling raw GPU time, it provides the tooling and managed infrastructure to run models as reliable, scalable services — a niche that grew rapidly once generative AI created mass demand for model serving.
The broader context is a maturing AI infrastructure market. The first phase of the boom concentrated capital on training compute and the data centers to house it. By 2026, attention had broadened to inference — the operational layer where AI meets users — drawing large financings to companies across the serving stack, from GPU clouds to optimization software.
Politico reported on June 18, 2026 that the Federal Energy Regulatory Commission (FERC) — characterized in the piece as “not the old sleepy agency” — is diving into the escalating fight over how data centers connect to the U.S. power grid. The report frames the once low-profile regulator as an increasingly active and decisive player in disputes over data-center interconnection, the process by which large new electricity loads are studied, approved, and physically wired into the grid.
Executive Summary
The headline itself is the story: a Washington energy regulator that historically operated far from public attention is now central to one of the most consequential infrastructure questions of the decade — how, where, and on what terms the data centers powering artificial intelligence get their electricity. Politico’s framing, that FERC is no longer “the old sleepy agency,” signals that the commission is taking an assertive posture in interconnection disputes rather than leaving them to utilities, regional grid operators, and states to sort out.
For the data-center industry, this matters because grid access — not land, capital, or chips — has become the binding constraint on new capacity in many U.S. markets. Whatever rules FERC shapes for connecting very large loads will influence project timelines, cost allocation, and site selection across the country. The report we are working from is a headline-level summary rather than a full text, so the specific proceedings, orders, or disputes Politico describes are not detailed here; our analysis focuses on why FERC’s posture matters and what remains to be confirmed.
Why the Grid Regulator Suddenly Matters to AI
FERC regulates interstate electricity transmission and wholesale power markets — the high-voltage backbone of the grid — and oversees the regional transmission organizations that run much of it. For decades that made it consequential mainly to utilities and power traders. The AI buildout changed the audience. Data centers are now proposing loads measured in the hundreds of megawatts and even gigawatts, on par with heavy industry or small cities, and connecting loads of that size raises exactly the questions FERC referees: who gets studied first, what upgrades are required, and who pays for them.
The “sleepy agency” framing in Politico’s headline captures a real shift in stakes. When interconnection was routine, the rules governing it were obscure. When interconnection becomes the gating item for a multi-hundred-billion-dollar industry, the same rules become front-page policy — and the body that writes them becomes a power broker whether it seeks the role or not.
The Interconnection Bottleneck Is the Business Story
Interconnection — the engineering and contractual process of plugging a new generator or large customer into the grid — has become notorious for multi-year queues in many U.S. regions. For data-center developers, an interconnection timeline is effectively a revenue timeline: a site that cannot energize cannot sell capacity. That is why disputes over queue rules, study procedures, and arrangements such as co-locating data centers directly at power plants (sometimes called behind-the-meter siting, where the load connects at the plant rather than through the wider grid) have turned into hard-fought regulatory battles.
How FERC resolves these fights will shape winners and losers. Clear, faster federal rules would favor developers with strong utility relationships and sites near existing capacity. Restrictive or unsettled rules push projects toward states and utilities perceived as easier to work with, toward on-site generation, or toward markets abroad. Utilities and existing ratepayers, meanwhile, have a direct stake in ensuring that grid upgrades driven by data-center demand are paid for by the companies that cause them rather than spread across household bills — a cost-allocation question that sits squarely in FERC’s lane.
An Assertive FERC Cuts Both Ways
An engaged regulator is not automatically good or bad news for the industry. On one hand, federal clarity could standardize how very large loads are treated, reducing the state-by-state and utility-by-utility uncertainty that currently complicates siting decisions. On the other, active federal scrutiny can slow novel deal structures — such as dedicated supply arrangements between power plants and data centers — while the commission works out reliability and fairness implications for everyone else on the grid.
It is also worth noting what FERC does not control. Siting of the data centers themselves, retail electricity rates, and most generation permitting remain state matters. So even a maximally assertive FERC is one decisive player among several, and the practical outcome for any given project will depend on how federal interconnection policy interacts with state regulation and utility planning. The Politico headline tells us the referee has taken the field; the source available to us does not detail which specific calls it is making.
Background
FERC traces its lineage to the Federal Power Commission, created in 1920, and has long operated as a technical regulator of interstate power transmission, wholesale electricity markets, and natural-gas infrastructure. Its rules govern the regional transmission organizations — such as PJM in the mid-Atlantic — that manage the grid across much of the country, and its interconnection procedures determine how new generators and, increasingly, very large customers plug in.
The agency’s rising profile tracks the AI-driven surge in electricity demand. After roughly two decades of flat U.S. power consumption, forecasts turned sharply upward in the mid-2020s as hyperscale data centers multiplied, and disputes over connecting them — including high-profile fights over siting data centers directly at power plants — began landing at FERC’s door. The June 2026 Politico report captures the resulting role reversal: an agency once known mainly to energy lawyers is now a decisive venue for the infrastructure economics of AI.
The Public Utility Commission of Texas (PUCT) has finalized new standards governing how large data centers connect to, and operate on, the state’s power grid, Houston Public Media reported on June 17, 2026. The rules implement Senate Bill 6, the 2025 Texas law that created a distinct regulatory category for very large electricity users — including data centers — seeking to plug into the ERCOT grid.
The action makes Texas the first U.S. state to complete a comprehensive rulebook for large-load interconnection and emergency curtailment at a moment when AI-driven data center demand is reshaping utility planning nationwide.
Executive Summary
Texas regulators have closed the loop on a process that began with Senate Bill 6, signed into law in June 2025. That statute directed the PUCT and ERCOT — the Electric Reliability Council of Texas, which operates the grid serving roughly 90 percent of the state’s electric load — to build new rules for “large loads,” generally facilities demanding 75 megawatts or more. The law’s core provisions required large customers to share better information during interconnection studies, bear more of the study costs, and accept that the grid operator can curtail (temporarily reduce or disconnect) their power during genuine grid emergencies.
Why it matters: Texas hosts one of the largest and fastest-growing data center pipelines in the world, and ERCOT’s interconnection queue has swelled with speculative large-load requests that make demand forecasting difficult. Finalized standards convert a statutory framework into operational reality — telling developers what they must disclose, what they will pay, and under what conditions their megawatts can be interrupted.
Because Texas is both the most active battleground for AI infrastructure siting and an energy-only market that other regions watch closely, these standards are widely expected to serve as a template. Utilities and regulators in other high-growth markets face the same problem Texas confronted first: how to welcome enormous new loads without socializing their costs or risking reliability for everyone else.
Why Texas Moved First
ERCOT operates an electrically isolated grid with limited connections to neighboring systems, which means Texas cannot import its way out of a supply crunch. When data center developers began filing interconnection requests at unprecedented scale, the gap between requested capacity and capacity that will actually be built became a planning hazard: transmission gets sized, and costs get allocated, against demand that may never materialize. Senate Bill 6 was the legislature’s answer, and the PUCT’s finalized standards are the machinery that makes it enforceable.
The economics are straightforward. Interconnection studies, transmission upgrades, and reserve capacity all cost money. Without rules assigning those costs to the large loads that trigger them, they flow to ordinary ratepayers. Texas has effectively decided that hyperscale demand should arrive with obligations attached — better data, upfront fees, and flexibility during emergencies — rather than as an unconditional guest.
Curtailment Changes Data Center Math
Curtailment — the grid operator’s ability to reduce or interrupt a customer’s power draw during scarcity events — is the provision with the sharpest commercial edge. Data centers sell uptime; their customer contracts are built on availability guarantees measured in fractions of a percent. A regulatory regime in which ERCOT can order large loads offline during firm load shed events forces operators to invest in the mitigations SB 6 contemplated: on-site backup generation, batteries, and workload orchestration that can shift compute out of state during grid stress.
That is not necessarily bad news for the industry. Facilities that can flex have something to sell — demand response is compensated in ERCOT — and AI training workloads, unlike real-time transaction processing, can often tolerate interruption. The standards effectively reward operators who engineer for flexibility and penalize those who assumed firm power was an entitlement. Expect the gap between those two designs to show up in siting decisions and financing terms.
A Template Other Grids Will Copy
Regulators in other high-growth markets — Virginia, Georgia, Arizona, and the multi-state PJM region — are wrestling with the same questions Texas has now answered on paper: who pays for network upgrades, how to filter speculative interconnection requests, and whether the largest loads should be interruptible. A finalized Texas rulebook gives them working language and, in time, empirical results to point to.
The competitive question is whether the standards make Texas more or less attractive. Developers may bristle at curtailment exposure, but regulatory certainty has value: a known process with known costs can beat a friendlier jurisdiction where interconnection timelines are unbounded. If Texas continues to land marquee AI projects under these rules, the argument that clear obligations deter investment will weaken, and the template will spread faster.
Background
Texas has become one of the world’s most important data center markets, drawn by cheap land, fast permitting, abundant natural gas and renewable generation, and an energy-only electricity market. That growth accelerated dramatically with the AI buildout, pushing ERCOT’s long-term demand forecasts sharply upward and filling its interconnection queue with large-load requests whose eventual construction was far from certain.
Senate Bill 6, passed by the Texas Legislature and signed in June 2025, was the state’s structural response: it required large electricity users to disclose more information, shoulder interconnection study costs, and accept curtailment authority during grid emergencies, then directed the PUCT to write implementing rules. The standards finalized in June 2026 are the culmination of that rulemaking.
A California water utility is investigating a claim by an Iran-linked threat actor that it breached the utility’s systems, according to a June 17, 2026 report from Cybersecurity Dive. As of the report, the intrusion is a claim under investigation — not a confirmed compromise — and the utility has not publicly validated the actor’s assertions.
Executive Summary
The report is short on confirmed detail but long on significance: a threat actor publicly associated with Iran has asserted that it compromised a water utility in California, and the utility has opened an inquiry into whether the claim is real. In critical-infrastructure security, that sequence — public breach claim first, verification later — has become a recurring pattern, and it matters regardless of how the investigation resolves.
Water and wastewater systems sit at the intersection of two uncomfortable facts. They are unambiguously critical infrastructure — a service failure has immediate public-health consequences — and they are, as a sector, among the least-resourced operators of industrial control technology in the United States. That combination makes them attractive targets for state-aligned actors seeking psychological and political impact, whether or not a given claim reflects a genuine operational compromise. For operators of data centers, networks, and other critical facilities, the episode is a reminder that adversary messaging is itself part of the attack, and that the ability to rapidly verify or refute a breach claim is now an operational capability in its own right.
A Claim Is Not a Breach — and That Distinction Is the Story
Everything public in this report hinges on the word “probes.” The utility is investigating; it has not confirmed an intrusion, and the actor’s assertion stands unverified. That matters because state-aligned and hacktivist-branded groups have a documented history of exaggerating, recycling, or fabricating claims against high-visibility targets. Publicly claiming a water-system breach generates headlines and anxiety at essentially zero cost to the attacker, whether or not any system was touched.
At the same time, dismissing such claims outright would be equally unwarranted. Iranian-affiliated actors have previously carried out real, confirmed intrusions against U.S. water utilities — most visibly the late-2023 wave of attacks on internet-exposed Unitronics programmable logic controllers, which defaced operator screens at multiple utilities and prompted advisories from CISA and the water sector’s information-sharing bodies. The honest posture, for readers and for the utility itself, is disciplined agnosticism: treat the claim as unproven, investigate as if it could be true, and communicate what is and is not known.
Why Water Utilities Keep Appearing in the Crosshairs
Water systems run on operational technology, or OT — the industrial controllers, sensors, and SCADA (supervisory control and data acquisition) software that open valves, run pumps, and dose chemicals. Much of this equipment was designed decades ago for reliability, not for exposure to a hostile internet, and many of the roughly 50,000 community water systems in the U.S. are small operations without dedicated cybersecurity staff. Remote-access tools bolted on for operator convenience, default credentials, and flat networks between office IT and plant floors are recurring findings across the sector.
For a state-aligned actor, this asymmetry is the appeal. Even a shallow intrusion — a defaced control screen, exfiltrated documents, a screenshot of an operator interface — can be presented as evidence of reach into an adversary nation’s drinking water, with psychological effect far exceeding the technical sophistication involved. The attacker’s goal is often the announcement as much as the access. That is why federal agencies have repeatedly urged water utilities to remove control systems from the public internet, enforce multifactor authentication, and change default passwords: measures that are basic, but that close precisely the doors these campaigns walk through.
The Verification Problem Is Now an Operational Cost
When a breach claim surfaces publicly, the target inherits an urgent, expensive burden: prove or disprove it, fast, under public scrutiny. That requires log retention deep enough to reconstruct weeks or months of access, asset inventories accurate enough to know what “our systems” even means, and forensic readiness in OT environments where taking a controller offline for imaging can interrupt service. Utilities that lack these capabilities face prolonged uncertainty — and prolonged uncertainty, not the intrusion itself, often does the most reputational damage.
There is a broader lesson here for every critical-infrastructure operator, including the data-center and connectivity industry. Incident response planning has traditionally started at detection; it increasingly needs to start at allegation. The ability to say, credibly and quickly, “we have investigated and here is what we found” depends on investments made long before any claim appears — monitoring of OT networks, segmentation between IT and control systems, and rehearsed communication plans. Those investments are unglamorous, but this episode shows exactly when they pay off.
Background
The U.S. water sector comprises tens of thousands of mostly small, locally governed utilities, and it has repeatedly been flagged by federal agencies as a cybersecurity soft spot among the sixteen designated critical-infrastructure sectors. Unlike bulk electric power, water has no binding federal cybersecurity standards regime of comparable reach, leaving practices uneven across systems of very different sizes and budgets. Iranian-affiliated threat activity against the sector is not hypothetical: the 2023 compromises of Unitronics control devices at several U.S. utilities — carried out by actors the U.S. government linked to Iran’s Islamic Revolutionary Guard Corps — demonstrated that opportunistic attacks on exposed water-system equipment do occur, and prompted sector-wide advisories on securing internet-facing controllers. Against that history, public breach claims aimed at water utilities land on well-prepared soil, which is precisely why each new claim demands careful verification rather than reflexive acceptance or dismissal.
CoreWeave, the GPU cloud provider, announced on June 17, 2026 that Kimi K2.7 Code — a coding-focused model in Moonshot AI’s open-weight Kimi family — is now available on its serverless inference service. The company says the offering delivers leading benchmark price-performance, positioning it as a low-cost way to run one of the more capable open coding models without managing GPU infrastructure.
Executive Summary
The announcement itself is narrow: a new model added to an existing managed service. Its significance lies in what it represents. CoreWeave built its business renting raw GPU capacity to AI labs and enterprises; serverless inference — where customers pay per token processed rather than per GPU-hour — is a move up the stack into a managed service business with different economics and a much broader addressable market.
The choice of model is equally telling. Coding models are among the most token-hungry workloads in AI today, because autonomous coding agents read and write large volumes of text in long loops. By pairing a well-regarded open-weight coding model with a price-performance pitch, CoreWeave is targeting exactly the segment — developer tools and agentic coding platforms — where inference bills are growing fastest and buyers are most price-sensitive.
What the release headline does not settle is the substance behind the claim: the syndicated summary does not include the actual per-token pricing, the benchmarks cited, or the rivals compared against. The claim is plausible given CoreWeave’s infrastructure scale, but as published it is a marketing assertion awaiting verification.
GPU Clouds Are Climbing the Stack
CoreWeave’s core product has historically been infrastructure: large clusters of Nvidia GPUs leased to customers who bring their own software. Serverless inference inverts that model. The provider runs the model, handles scaling and reliability, and bills per token — the unit of text an AI model reads or writes. For customers, this removes the hardest parts of AI operations: capacity planning, GPU utilization, and model serving expertise.
For CoreWeave, the strategic logic is margin and market breadth. Raw GPU rental is increasingly commoditized and dominated by a small number of very large contracts. A token-metered service can serve thousands of smaller customers, smooth utilization across its fleet, and capture software-layer value on top of hardware it already operates. Every major GPU cloud is attempting the same climb, which is precisely why price-performance has become the battleground.
Open-Weight Models Fuel an Inference Price War
Kimi K2.7 Code is part of Moonshot AI’s Kimi line of open-weight models — models whose trained parameters are published for anyone to download and run, unlike closed models such as those from OpenAI or Anthropic, which are available only through their makers’ APIs. Open weights turn model serving into a competitive market: many providers can host the identical model, so they compete on price, speed, and reliability rather than exclusive access.
That dynamic is good for buyers and brutal for margins. When the model is a commodity, the winner is whoever runs it most efficiently — better hardware utilization, better serving software, cheaper power. CoreWeave’s implicit argument is that owning and operating its own large-scale GPU fleet lets it undercut resellers and match or beat specialist inference providers. The claim is credible in principle; whether it holds depends on numbers the announcement headline does not supply.
Coding Is the Beachhead Workload
The decision to lead with a coding model is not incidental. AI coding assistants and autonomous coding agents consume tokens at rates far beyond chat applications, because they iterate: reading codebases, generating changes, running checks, and revising, often for many cycles per task. For the companies building those tools, inference cost is a first-order line item, and many of them already prefer open-weight models specifically so they can shop across hosts.
Winning this segment matters beyond the immediate revenue. Developer-tool companies are sophisticated, benchmark-driven buyers; a provider that earns their workloads gains both a proof point and a durable base of high-volume usage. Conversely, they are also the quickest to leave when a competitor posts a better price-per-benchmark-point, which keeps pressure on every provider’s pricing.
Reading Price-Performance Claims Carefully
“Leading benchmark price-performance” is a compound claim, and each half deserves scrutiny — as it would from any vendor. On the performance side, coding benchmarks are useful but imperfect proxies; results can vary with how a model is configured and served, so a hosted version’s scores should ideally be verified against the model publisher’s own reported figures. On the price side, headline per-token rates can obscure differences in speed, rate limits, context-length pricing, and reliability guarantees that materially change real-world cost.
None of this means the claim is wrong. It means the appropriate response, for any buyer, is a straightforward evaluation: run your own workload, measure quality and latency, and compute cost per completed task rather than cost per token. That standard applies equally to CoreWeave and to every competitor making similar claims in what has become a loudly contested market.
Background
CoreWeave rose from cryptocurrency-mining origins to become one of the most prominent specialized GPU clouds of the AI boom, operating large fleets of Nvidia accelerators for AI labs and enterprises, and completed its Nasdaq IPO in March 2025. Like other GPU clouds, it has been expanding from raw infrastructure into managed services — of which serverless inference is the most direct bid for the application-developer market.
Moonshot AI’s Kimi K2 family established itself as one of the leading open-weight model lines, drawing attention especially for coding and agentic tasks. Because the weights are published, the models are served by many competing providers worldwide — a dynamic that has made hosted open-weight inference one of the most price-competitive corners of the AI market, and the arena in which CoreWeave’s announcement stakes its claim.