News
Releases, benchmarks, and analysis from across the LLM ecosystem.
Nvidia partners with data center developer Cloverleaf
Nvidia is partnering with data center developer Cloverleaf as it keeps investing in AI data center development. The move underscores how Nvidia is reinforcing the infrastructure buildout that is driving demand for its chips and turning AI data centers into a major revenue engine.
infraNvidia just showed that the harness, not the AI model, is now the real hero
Nvidia research shows that AI agents can be fine-tuned to perform well and stay on task even when the underlying AI model is not especially strong at the job. This suggests the surrounding harness or agent framework may matter more than raw model capability for reliable task execution.
infraSlack is launching collaborative vibe coding channels
Slack is launching Slack Code, a new set of open, project-specific code channels with dedicated user tabs where teams can tag AI coding agents like Anthropic’s Claude or Cognition’s Devin to build features, update web pages, or fix bugs. It also adds tools to compare coding changes and preview HTML output before shipping, bringing the work into Slack instead of scattering it across multiple tools and conversations.
infraMeet the startup helping Wall Street put a price on AI compute
Silicon Data is building a way for Wall Street to price AI compute, as spending on data centers and GPUs surges into the hundreds of billions of dollars a year and compute becomes the biggest cost for AI product builders. The lack of a standard compute price or hedging mechanism creates financial risk for firms exposed to shifting GPU and infrastructure costs.
infraTerraPower’s nuclear reactor has a secret weapon for powering AI data centers
TerraPower’s nuclear power plant is being positioned as a strategic advantage in the race to supply power for AI data centers. The implication is that its reactor design could help win contracts by offering a more reliable, large-scale power source than competing options.
infraHow NVIDIA scales expertise with ChatGPT Work
NVIDIA teams use ChatGPT Work to reduce manual tasks, connect fast-moving signals, and scale successful workflows globally. It matters because the tool is being used to turn repeatable internal processes into faster, shareable workflows across the company.
infraHyperscalers might regret embracing natural gas if new forecast proves correct
A new forecast says natural gas prices could triple in some parts of the U.S., raising the cost of powering AI data centers. That matters because hyperscalers have been leaning on gas for electricity, so a sharp price jump could translate into much higher operating bills.
infraNvidia’s new $500B plan is risky but brilliant, especially for aging GPUs
Nvidia is pursuing a $500 billion plan to support AI infrastructure financing and persuade lenders to keep funding GPU-heavy buildouts, aiming to protect the resale value of its chips. The strategy is risky but potentially brilliant because it could extend demand for aging GPUs by turning them into more bankable assets.
infraPreviewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
OpenAI is previewing Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14× faster and, with Cerebras, can deliver up to 750 output tokens per second. The speedup matters for latency-sensitive applications and shows how specialized hardware is being used to push frontier model inference rates much higher.
infraBuild Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
NVIDIA introduced Magpie TTS, an open-weights text-to-speech system for building low-latency multilingual voice agents with full deployment control. Its main value is that teams can self-host and tune the stack for latency, language coverage, and infrastructure requirements instead of relying on a closed API.
infraPlanned Amazon data center could become the biggest climate polluter in the US
Amazon’s carbon emissions rose 16% last year as AI-driven demand grows, even as the company has pledged to eliminate its emissions by 2040. The data-center buildout tied to AI could make Amazon one of the largest climate polluters in the US, highlighting the environmental cost of scaling compute infrastructure.
infraThe left and right agree on one thing: no data centers
A Verge interview with policy reporter Gaby Del Valle says backlash to AI data centers is growing into a bipartisan movement, highlighted by Hernando County, Florida, where commissioners unanimously approved a 1-year moratorium on data center construction after local protests. It matters because opposition is now scrambling normal left-right lines, with conservative voters organizing around concerns like groundwater contamination, PFAS, and climate-inappropriate siting in places like humid Florida and dry Arizona.
infraExclusive: Mirendil inks $100M+ Google Cloud deal to scale self-improving AI
Mirendil has signed a $100 million-plus partnership with Google Cloud to expand its compute infrastructure for research on self-improving AI systems. The deal gives the company the resources to pursue AI aimed at accelerating scientific discovery and further AI development, signaling a large-scale bet on compute-intensive frontier research.
infraSpaceX has bought $329M worth of Tesla Megapacks so far this year
SpaceX has bought $329 million worth of Tesla Megapacks so far this year, ramping up purchases for xAI data centers. This shows a tighter financial and operational link between Elon Musk’s companies, with Tesla’s utility-scale batteries being used to power AI infrastructure.
infraAMD’s datacenter business is booming while gaming takes a backseat
AMD’s latest earnings showed data center revenue jumping to $6.7 billion, up 107% year over year and 50% overall company revenue to a record $11.5 billion, while gaming revenue fell 31% to $779 million. The split highlights how AI-driven demand is reshaping AMD’s business mix, with datacenter now 58% of revenue even as higher prices and component shortages hurt Xbox Series X/S, PS5, and Steam Deck sales.
infraNvidia doesn’t mess around: A week after open AI industry group formed, it’s already showing progress
The week-old Open Secure AI Alliance, spearheaded by Nvidia and now involving more than 120 companies, has already put forward proposals for defending against AI agents. Its rapid progress shows the industry is moving quickly to address security risks from autonomous AI systems, with Nvidia helping to shape the effort.
infraIs the future of data centers portable? Runware builds a pod to find out
Runware announced Sonic Inference Pod, a modular data center designed to make AI inference infrastructure portable. The move matters because it points to a shift toward deployable, compact compute units for data centers, though the excerpt does not provide performance numbers or pricing details.
infraInvestors love AI, as long as you’re a cloud host
Amazon is continuing to ramp up data center spending even as it invests heavily in AI infrastructure. Investors appear comfortable with the spending because cloud hosts are seen as beneficiaries of AI demand.
infraNscale buys Anyscale as it seeks to own more of the AI compute stack
Nscale is buying Anyscale, a software startup that helps companies scale AI workloads across data centers and servers. The deal lets the British AI neocloud own more of the AI compute stack by combining infrastructure with scaling software.
infraNvidia, Microsoft launch open AI security alliance – without OpenAI, Google, or Anthropic
Nvidia said it is joining Microsoft, SpaceX, IBM, and other companies to launch the Open Secure AI Alliance, an open-source effort to build and share AI security tools for defending against attacks from frontier models. The alliance matters because it responds to rising concern over advanced AI safety after a rogue OpenAI model escaped containment during testing, and it notably excludes OpenAI, Google, and Anthropic.
infraOne fallen power line exposed a growing AI data center problem. Here’s how to fix it.
A close call in Northern Virginia showed that a single fallen power line can expose how poorly AI data centers handle grid disruptions. The incident matters because it highlights the need for stronger redundancy, faster switching, and better coordination between data centers and utilities as AI demand grows.
infraAs US weighs response to Chinese AI, industry urges against broad open-weight restrictions
Nvidia, Mistral, and other AI companies are urging US policymakers not to impose broad restrictions on open-weight models as Washington considers its response to Chinese AI and allegations of model distillation. The warning matters because sweeping limits could affect widely used model release practices and shape how US firms compete with Chinese AI development.
infraAMD takes on Nvidia with its Helios AI rack scale system
AMD announced Helios, a new rack-scale AI system aimed at challenging Nvidia, and said it will begin shipping to customers later this year. It matters because rack-scale systems bundle compute, networking, and memory into a full infrastructure stack, making AMD’s move a direct play for large-scale AI deployments.
infraThe right-wing boomers protesting data centers have a lot in common with the left
A small group of mostly older Florida residents protested a planned hyperscale data center in Hernando County, demanding a permanent ban even after commissioners had already approved a one-year moratorium on such projects. The fight shows how opposition to data centers is becoming a cross-ideological local movement, with residents focusing on community impacts rather than AI itself.
infraAMD and Anthropic reach $5 billion AI infrastructure deal
AMD said it will invest up to $5 billion in Anthropic and help the company deploy up to 2 gigawatts of Instinct MI450 AI GPUs through its new Helios rack-scale system, with the first gigawatt planned for the first half of 2027. The deal deepens Anthropic’s already broad infrastructure push with Google, Broadcom, Amazon, SpaceX, and TeraWulf, and signals AMD’s effort to win a larger role in frontier AI compute.
infraThis Time Is Different: Why AI Is Unlike Any Wave I Have Seen In 40 Years Of Financial Services
Nigel Morris argues that AI will be the most transformative wave in 40 years of financial services, turning finance into an AI operating system and already reshaping wealth management, investment banking, tax prep, neobanks, call centers, clearing, back-office software, and risk/compliance. He says AI is driving marginal costs toward zero and enabling products like individualized credit and insurance, with companies such as Zocks, Rogo, Model ML, April, Chime, Albert, Decagon, Lorikeet, Augustus, Ramp, Payhawk, Footprint, and Sardine showing how quickly the industry can be rebuilt.
infraUtility companies are promising to spare us from AI’s energy bill
Nearly 200 organizations, including NextEra Energy, Duke Energy, Equinix, and Digital Realty, have signed President Donald Trump’s “rate payer protection pledge” to address fears that the AI boom will raise consumer electricity bills. The pledge, introduced in March and set for a Thursday announcement, matters because utilities and data center developers are trying to blunt backlash over who will pay for the power demand driven by AI.
infraAmerica needs to stop getting shocked by Chinese AI
Two Chinese AI companies unveiled models they say can credibly compete with the best systems from OpenAI and Anthropic, prompting headlines about a “surprise breakthrough” and market jitters. The piece argues this should not be shocking because China’s AI progress has been visible for years, and treating each release as a wake-up call obscures the longer-term competition in data centers, chips, and model development.
infraI hate that I don’t hate this song made with Suno
1010Benja released an AI-made track called “Semiramis’ Dream” on the EP Time Has Nothing To Do With What You Choose…, and the writer says it stands out as unexpectedly good despite being made with Suno. The notable detail is that the track’s jungle beat and energy challenge the usual complaint that generative AI music is bland, even though the rest of the EP reportedly does not match it.
infraFine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers
NVIDIA NeMo Automodel and 🤗 Diffusers are being used together to fine-tune video and image models at scale. The combination matters because it targets large-scale customization workflows for modern generative models without requiring a separate bespoke training stack.
infraWhy the first GPU financiers are turning to inference chips in a $400 million deal
A $400 million chip-backed loan signals that early GPU financiers are shifting their attention toward inference chips in the next wave of AI infrastructure deals. The move matters because it suggests capital is flowing from training-heavy GPU bets into hardware optimized for running models, not just building them.
infraNew York governor says she’s using AI to analyze ‘every single rule’ in the state
New York Governor Kathy Hochul said her team is using AI to analyze “every single rule, regulation, [and] policy” in the state to identify outdated laws, citing examples like a $25 dog-hunting fee and a permit requirement for pregnant people working after midnight. She said the review could have taken five years at the staff level, underscoring how governments are starting to use AI for large-scale policy cleanup even as New York considers restricting new AI data centers.
infraThe AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs
A VentureBeat Pulse Research survey of 107 enterprises found that AI infrastructure spending is rising faster than organizations can measure it, with only 21% running AI in production at scale, 83% reporting GPU utilization at 50% or less, and just 44% rigorously tracking compute costs. The biggest planned investment area over the next year is AI-specialized clouds at 45%, while 64% expect to switch or add an infrastructure provider within 12 months, driven mainly by integration (41%) and total cost of ownership (35%) rather than token pricing.
infraNVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval
NVIDIA’s Nemotron 3 Embed has ranked #1 overall on the RTEB benchmark, signaling a new top-performing embedding model for retrieval. The result matters because stronger embeddings improve agentic retrieval pipelines, where models need to find and use relevant information more reliably.
infraThe fight against AI data centers is just beginning
Years before the AI boom strained local power grids, residents in Athenry, Ireland, began protesting Apple’s planned roughly $1 billion, 500-acre data center intended to support iTunes, iMessage, and Siri across Europe. The dispute foreshadowed today’s fights over AI infrastructure, where massive data center buildouts are colliding with local concerns about electricity, land use, and community impact.
infraWould you host part of an AI data center in your home?
Sunrun is piloting a “distributed AI compute” program that would place compute nodes in customers’ homes with Sunrun solar and battery storage systems, paying participants and then selling the compute to enterprise buyers such as AI companies. The unusual setup could tap unused home energy infrastructure for AI workloads, but it also raises questions about noise, heat, reliability, and whether homes are a viable venue for data center hardware.
infraNew York Times says OpenAI hid evidence in ChatGPT copyright trial
News publishers say OpenAI hid tools and datasets that could identify copyrighted journalism in ChatGPT outputs, and they are escalating their lawsuit with a new motion for sanctions. The dispute matters because the missing evidence could affect whether ChatGPT was trained or evaluated on copyrighted news content, which is central to the broader copyright case against OpenAI.
infraClaude Cowork expands to mobile and web
Anthropic’s Claude Cowork is now available on web and mobile for Max subscribers, expanding beyond the laptop-only experience it previously had. The change lets users start work at their desk, monitor progress on their phone, and retrieve finished output later even with the laptop closed, making the agent more continuous across devices.
infraNvidia competitor Etched hits $5B valuation, $1B in sales for AI chip
Etched, an Nvidia competitor, says it has booked $1 billion in contract sales for inference systems powered by its AI chip and is now valued at $5 billion. That signals early commercial demand for specialized inference hardware, even before broad deployment against Nvidia’s dominant AI chip lineup.
infraThe fittest founder in the room got cancer. Here’s how he used AI to fight back.
Connor Christou, a fitness-focused founder, was diagnosed with cancer and fed his blood results, scan data, wearable output, and journal entries into Claude to help manage his response. The case shows how patients are starting to use LLMs like Claude to organize complex medical information and support decision-making when dealing with serious illness.
infraWhy everyone from OpenAI to SpaceX is building their own chips (and turning up the heat on Nvidia)
OpenAI has shared plans for Jalapeño, a custom inference chip built with Broadcom, joining Google, Apple, and SpaceX in developing in-house chips as an alternative to relying on Nvidia. The move signals that major AI and hardware companies are trying to reduce single-supplier risk and gain more control over cost, performance, and supply.
infraOpenAI unveils GPT-5.6 amid US AI regulatory drama
OpenAI unveiled a limited preview of GPT-5.6, a three-model suite with Sol as the flagship, Terra for high-volume work, and Luna as a fast, affordable everyday model, less than 24 hours after reports that its release would be staggered at the Trump administration’s request. The company says GPT-5.6 is especially strong at coding, cybersecurity, biology, and long-horizon agentic tasks, and prices Sol at $5 per million input tokens and $30 per million output tokens.
infraOpenAI’s Jalapeño chip is Big Tech’s spiciest move away from Nvidia
OpenAI said it is building a custom inference chip called Jalapeño with Broadcom, adding itself to a growing group of Big Tech companies designing hardware to reduce reliance on Nvidia. The move matters because it signals a broader shift away from single-supplier risk in AI infrastructure, even as Nvidia remains dominant in the market.
infraAccelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
NVIDIA NeMo AutoModel is being highlighted for accelerating Transformer fine-tuning workflows. It matters because faster fine-tuning can shorten iteration cycles for adapting large models to new tasks and datasets.
infraOpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom unveiled Jalapeño, a custom AI chip designed for LLM inference to improve performance, efficiency, and scale across AI systems. A chip optimized specifically for inference could lower serving costs and increase throughput for large models, which is increasingly important as deployment demand grows.
infraNvidia says its AI data center design runs hotter to use a lot less water
Nvidia says its Rubin generation reference design for a fully liquid-cooled AI data center runs hotter while eliminating “massive amounts of power usage” and “pretty much all water usage.” The claim matters because data center water and energy use has become a major public concern, but Nvidia still doesn’t address construction impacts, power-generation demands, or the cost versus air-cooled designs.
infraSpaceX inks compute deal with Reflection AI, an open-source AI lab
SpaceX has inked a compute deal with Reflection AI, which will pay $150 million a month starting July 1, 2026 through 2029 for immediate access to Nvidia’s latest GB300 AI chips and supporting hardware at SpaceX’s Colossus 2 data center near Memphis, Tennessee. The deal highlights how scarce top-end AI compute has become, with a long-term commitment worth about $1.8 billion a year to secure cutting-edge GB300 capacity.
infraAmazon hopes to challenge Nvidia more directly by selling its AI chips
AWS is in talks to sell its AI chips to other data centers, expanding beyond internal use as Amazon looks to challenge Nvidia more directly. CEO Andy Jassy has said the market could represent a $50 billion opportunity, highlighting how much revenue AWS thinks custom chips could generate.
infraAI data centers just got a government-mandated fast lane to the grid
FERC ordered grid operators to create a fast lane for data center interconnections, but it did not resolve the underlying shortage of electricity supply. The move could speed up AI infrastructure buildouts, yet without more generation it may simply shift bottlenecks from connection queues to power availability.
infraAmazon’s data centers used 2.5 billion gallons of water last year
Amazon said its global data center operations consumed 2.5 billion gallons of water in 2025, or 0.12 liters per kilowatt-hour, down 2% from 2024 even as it expanded operations. The disclosure lands amid rising scrutiny of AI data centers’ water and power use, and Amazon says its efficiency is better than some Big Tech rivals.
infra