SUMMARYRidley Scott is developing The Dog Stars, a postapocalyptic film set after an aggressive flu devastates much of the population, and he is also working on a Treasure Island adaptation starring Hugh Jackman. He has begun rewriting a third Alien prequel and plans to continue exploring the fallout from a malfunctioning AI character. Scott says he still relies on hand-drawn storyboards and has not yet adopted AI tools in his filmmaking.

The Los Angeles Times first calls Ridley Scott's newest movie "a postapocalyptic thriller that trades zombies for pandemics and asks whether hope can outlast catastrophe." But besides The Dog Stars, Scott is also working on an adaptation of Robert Louis Stevenson's "Treasure Island" starring Hugh Jackman, and is thinking about turning one of his past films into a musical. ("I can't tell you which one because they're being difficult.") The common theme, they suggest, is a director who allows his stories to "speak through evocative, immersive images rather than unnecessary exposition."

The British director says it's because he got his start in advertising after attending the Royal College of Art in London. After 1982's neon-coated sci-fi thriller "Blade Runner," Scott notably made the iconic Apple Macintosh computer ad "1984," a memorable Orwellian-themed clip that contained suggestions of the dystopias that would continue to fascinate Scott in his filmmaking. "Advertising in England at that point became a competitive art form," he says, pointing to the notorious "Adland Five," a group of directors that included Hugh Hudson, Adrian Lyne, Alan Parker, Scott and his brother Tony. "I've got a show reel which is 50 years old which would entertain the s - out you. Hasn't aged at all. Our visual drive changed the face of how films looked." That's apparent in Scott's latest, "The Dog Stars" (in theaters Friday), a film about survival after the end of the world. It's based on Peter Heller's 2012 novel and set in Colorado 10 years in the future after an aggressive flu has decimated much of the population. Those remaining alive use any means necessary to stay that way, resulting in instability, brutality and, more hopefully, connection... These days, Scott develops movies with his production company, Scott Free, which he founded with his late brother Tony in 1995. Before that he says he struggled to find good scripts. And rarely has someone handed one to him. "I've only ever had three films land on my desk," he says. "One was 'Alien,' which I was fifth choice for after Robert Altman." He makes a face. "Can you imagine offering Robert Altman 'Alien'? How stupid can you get?" Another was 2015 space adventure "The Martian," which Scott says had been sitting on a shelf for two years before it was sent to him. Scott's grounded, semi-comedic take on Andy Weir's novel resulted in seven Oscar nominations. The third screenplay to materialize was "The Dog Stars." "I read it and I was blown away," Scott says.

Deadline notes that Scott is already re-writing the third prequel to Alien, citing Scott's comments to French outlet AlloCiné.

"The one that just came out [2024's Alien: Romulus] was OK. I think it needs some help. So, I've gone back in." Scott continued, "I've already got a footprint for the next Alien.... I'm picking up where we left off in Covenant with Micheal Fassbender, who now thinks he's Ozymandius." With the sequel bringing back Fassbender's futuristic AI droid, Scott sees an opportunity to explore the repercussions of artificial intelligence in our current world. "I like the fact that AI has gone off the rails," said Scott of where the character is at the end of Covenant. "If we go AI, the moment we discover it's off the rails is too late. It could switch every digital thing off in the world like that. You'd have fucking chaos. It's that kind of stuff that's scary."

Scott tells the Los Angeles Times that the malfunctioning AI character fascinates him, and in the movie will be creating a "brave new world for himself... AIs can go off. And if they do, it's like a bomb..."

So far, Scott says he hasn't implemented AI technologies into his filmmaking. He holds up his "Treasure Island" Ridleygrams [hand-drawn storyboards] and shakes them at the camera. "I can compete with an AI without any problem making images," he says. "But it's coming, inevitably." Scott was not involved in the forthcoming Prime Video series "Blade Runner 2099," out in November. "I can't blame my lawyers and my agents at that time, but somehow we let 'Blade Runner' go," he says. "And when you're not the author of the actual written word, you don't have a piece." (He was more involved in Noah Hawley's recent "Alien: Earth" series, which he calls "pretty damn good....")

Does that mean he thinks about his legacy as a filmmaker? He does, but not often. "When you think of your legacy, you're too near the end," Scott says. "I'm not near the end." He looks up from his drawings. "I've had a riotous time," he says. He corrects himself: "I'm having a riotous time."

SUMMARYScientists testing a CRISPR-Cas9 infusion found that a one-time gene edit lowered LDL cholesterol by about 50% and also reduced triglycerides in a tiny pilot study of 15 people with severe, medication-resistant cholesterol. Follow-up data published in November 2025 showed the effect lasted more than a year in the four participants who received the highest dose, with no adverse effects reported so far. If confirmed in larger trials, the approach could offer a durable treatment for people at high risk of heart disease.

CNN reports that "A one-time snip of a key gene lowered heart-clogging cholesterol by about 50% over an entire year with no adverse effects, scientists have reported."

The overwhelmingly positive results were found in four people who had taken the highest dose of the experimental treatment during a pilot study published in The New England Journal of Medicine in November 2025. That study was extremely small - only 15 patients with dangerously high cholesterol that did not respond to medication. The research was designed to test the safety of five different doses of a gene-altering infusion delivered by CRISPR-Cas9, a biological scissor which cuts a targeted genetic location to modify or turn a gene on or off. "If you'd asked me 15 years ago if we could have done something like this, I would have thought you were crazy," senior study author Dr. Steven Nissen, chief academic officer of the Sydell and Arnold Miller Family Heart, Vascular & Thoracic Institute at Cleveland Clinic in Ohio, told CNN previously.

An update on the study, published Friday in NEJM, appears to answer a key question: Does the intervention last? "Is something like this truly a one and done? That was always the question," said lead study author Dr. Luke Laffin, a preventive cardiologist at Cleveland Clinic's Heart, Vascular & Thoracic Institute. "This new data is essentially saying that the decreases in LDL cholesterol and triglycerides we saw at 60 days after treatment have lasted over a year," Laffin said. "In people who took the highest dose, the reduction seems durable and safe...."

If larger clinical trials show similar results, the procedure could be a game changer for young people with severe disease, preventive cardiologist Dr. Ann Marie Navar, an associate professor of cardiology at UT Southwestern Medical Center in Dallas, told CNN previously. "If you're 20 and you have really high cholesterol, it may make a lot more sense to have a one-time treatment that doesn't require you to have to take a pill every single day or shot every two weeks for the next 60 years," said Navar, who was not involved in the study. "The potential for this is just enormous."

Being born with this mutation naturally "drastically lowers or even eliminates a person's risk for heart disease," the article points out. And the study's senior author adds "now that CRISPR is here, we have the ability to change other people's genes so they too can have this protection."

SUMMARYRoblox is restricting games that reward younger players for watching endless autoplaying media feeds, a move aimed at experiences like Steal an Egg. The new policy blocks such games for Roblox Kids and Roblox Select accounts, while still allowing short rewarded video ads. Steal an Egg has also returned to the Roblox store without its video-feed feature.

Roblox is restricting games that reward younger players for endlessly consuming autoplaying media feeds. The new restrictions are specifically targeting Roblox experiences like "Steal an Egg," which encouraged players to watch TikTok-style clips for in-game benefits. "Roblox's official alert does not expressly mention SAE, but the new policy is most certainly targeting the viral game first and foremost," reports TechSpot. From the report: The company said games that reward ongoing media consumption are now unavailable to players with a Roblox Kids (ages 5 to 8) or Roblox Select (ages 9 to 15) account. The policy will affect games designed to combine media feeds, autoplay, infinite scrolling, and systems that reward players for doomscrolling through TikTok videos of questionable value.

"Steal an Egg" fits perfectly into this new category. The game tasks players with raising eggs, but they can also steal eggs from other NPCs. Players are incentivized to increase their characters' speed either by spending in-game currency (Robux) or by watching a video feed on a virtual treadmill. "Steal an Egg" was recently removed from Roblox but is now back in the store without its troublesome video-feed feature.

[...] Roblox also highlights that the new anti-doomscrolling restriction does not apply to the rewarded video ads feature, which rewards players for watching a short clip but does not include autoplay or other infinite-scroll-like design choices.

SUMMARYGoogle has added an Expert Intelligence feature to Gemini Notebook, formerly NotebookLM, that lets users cite more than 100,000 books they own through Google Play Books. The feature can turn book content into quizzes, infographics, and Audio Overviews, and Google is launching it first in the Gemini Notebook app and web dashboard before expanding it to the core Gemini app and Search AI mode. Google is also partnering with select authors on Featured Notebooks that add context beyond the books themselves.

Google is adding an "Expert Intelligence" feature to Gemini Notebook (formerly known as NotebookLM) that lets users cite and work directly from more than 100,000 books they own through Google Play Books. Those books can be used as sources for things like quizzes, infographics, and Audio Overviews. Digital Trends reports: For now, Google is offering one free book to users in the US as part of the launch campaign, with a caveat -- "while supplies last." At the moment, Expert Intelligence is only available in the Gemini Notebook app and its web dashboard, but down the road, Google says that it will also appear within the core Gemini app and AI mode in Search.

The broad idea behind expert intelligence is to combine what you can already do with Gemini, and expertise sourced from your favorite books. So, if you're asking Gemini to create a meal plan that is based entirely on chicken and spinach, you can ask the AI assistant to look for relevant recipes in one of Martha Stewart's cookbooks that you own. You can use these books as a source and turn them into Infographics, Audio Overviews, and Quizzes, among other formats. Google has partnered with select authors on Featured Notebooks that provide extra insights and context beyond their books. It's also worth noting that access is still tied to book ownership, so anyone using a shared Gemini notebook must own or purchase the book to fully use the feature.

SUMMARYVirtual power plants let utilities coordinate smart thermostats, EV chargers, home batteries, and solar panels to reduce electricity use during peak demand, often in exchange for bill credits or sign-up bonuses. More than 500 VPP programs operated in the US as of 2023, and millions of households are already enrolled, with interest growing from utilities and companies such as Google. The guide explains how to check eligibility, weigh flexibility and privacy tradeoffs, and decide whether the payments are worth it.

MIT Technology Review’s How To series helps you get things done.

Your thermostat may not look like a power plant. Neither does your electric vehicle, home battery, or HVAC system. But utility and energy companies increasingly want to treat them like one.

A virtual power plant, or VPP, is a collection of household devices (such as smart thermostats, electric-vehicle chargers, home batteries, and solar panels) that a utility can control. Usually that means commanding the devices to draw less electricity during peak hours. For example, the utility might adjust your thermostat or delay or slow EV charging when electricity demand is high.

In exchange, the utility offers VPP participants a discount on their energy bills and, in some cases, a signing bonus. Seth Frader-Thompson, CEO and cofounder of EnergyHub, a software company that helps utility companies run VPP programs, says a smart thermostat program may offer an initial bonus of roughly $50 to $150, plus about $25 to $50 per year, while home battery and EV devices could yield hundreds or thousands of dollars in annual savings.

The amount of power the utility might throttle in any one home is small. But it adds up, Frader-Thompson says. “When you put it together at the scale of hundreds of thousands, or millions, it has a pretty profound impact,” he says, equivalent to “firing up a power plant.”

As of 2023, there were already more than 500 VPP programs operating in the US alone, and the number has only grown since, especially with big players like Google starting to invest in this technology to help power their data centers. An estimated 4 million households with smart thermostats were enrolled in a VPP program as of last year.

But the approach is still new, and some programs may still have some kinks to work out, says Severin Borenstein, faculty director of UC Berkeley’s Energy Institute at Haas and member of the board of governors of the California Independent System Operator, which manages most of the state’s electric grid. If a program is not implemented well, he says, a utility may incorrectly predict when VPP participants plan to use more electricity and pay them for not using energy they weren’t planning to use anyway, potentially increasing energy bills for nonparticipants. Still, Borenstein says, “if we do it well, I think it can really be a benefit,” one that could help utilities avoid an expensive grid upgrade or emergency measures to conserve power.

Most consumer VPPs today are less dramatic than the name suggests and don’t actively send energy from your EV or home battery to the grid. But battery-to-grid programs are on the rise—and potentially offer even larger savings for consumers in the future.

So how do you actually sign up for a VPP? And how do you know if it’s worth it?

1. Check whether your utility company has a program and, if so, whether it actually supports your devices.

The types and brands of home devices supported vary from program to program. Your utility’s website is the obvious place to look to see if yours qualifies, but it’s important to note that you may not actually see the phrase “virtual power plant” anywhere. You may have better luck searching for your utility’s name plus terms like “demand response,” “peak rewards,” “connected solutions,” “battery storage,” “smart thermostat rewards,” “managed charging,” or “bring your own device.”

But don’t stop with the utility, Frader-Thompson says: “The way most people actually learn about this and sign up is through the manufacturer of the device they have.” In other words, the offer may show up through your smart thermostat app, EV app, or battery app, or in an email from the company that made the device.

Once you find a program, the instructions for enrollment may be as simple as clicking through an app, filling out a utility form, or confirming your account and device information through a third-party enrollment page. EV drivers may be able to see the terms and payment in their automaker app and enroll “with a click of a button,” says Joseph Vellone, CEO of the EV-focused VPP company ChargeScape.

Eligibility can get annoyingly specific. A smart thermostat program could require an approved Wi-Fi thermostat; an EV program may depend on your automaker, charger, utility territory, or rate plan; a battery program may depend on the battery brand, inverter, or installer and whether your system can communicate with the utility.

These programs are also not evenly distributed across the country. Most programs are established in places with lots of flexible devices, stressed grids, supportive utilities, or strong state policies—especially California, Texas, New England, and increasingly parts of the mid-Atlantic region.

2. Ask yourself how much flexibility you can afford.

Before you sign up for a VPP, you’ll want to determine whether you’re willing to let a company adjust a device in your home—even if it typically happens only a few times a week.

For some people, this may be an easy decision: If your EV sits plugged in all night but only needs two hours to charge, shifting when that charging happens may be almost invisible. A home battery program could be lucrative if you understand how often the battery will be used, how much backup power you can keep, and whether extra cycling affects your equipment.

Other households, however, “do not have the flexibility to engage in one of these programs,” says Sanya Carley, a professor at the University of Pennsylvania and faculty director of the Climate Center for Energy Policy. She says that people who work night shifts, have caregiving responsibilities or health needs, or are already aggressively limiting their energy use to save money may have less room to allow a utility to adjust heating, cooling, or charging rates during peak hours for grid demand.

3. Review the opt-out rules and read the fine print.

VPP programs generally give participants the ability to override temporary changes made by the utility. This right to “opt out” is what makes them workable for many customers. Can you skip a day of the program on your thermostat if you’re planning to have guests over? Can you tell your car to charge immediately before a long road trip? Can you keep a battery reserve for outages? Utilities are typically motivated to make the opt-out process as simple as possible, with few rules and restrictions.

It could also be worth investigating where your data might be going. EV and battery programs may need to collect data about things like charging status and schedule, or how much power a device is drawing, while smart thermostat data may reveal patterns about when people are home, sleeping, or using appliances.The Electronic Frontier Foundation, a nonprofit focused on digital rights, has warned that this data could be used to infer private routines inside a home; depending on the program, that information may not only move through a utility but get distributed to device manufacturers, software platforms, or third parties involved in running the program.

ChargeScape and Energy Hub say the data used for these programs is limited and functional. EV data is focused on “the physics and the energy of the asset itself,” Vellone says. Frader-Thompson explains,“It doesn’t really matter what any one customer is doing. It matters what the average customer is doing.”

4. Decide whether the offer is worth it for you.

The amount of compensation for signing up for a VPP can vary widely. The payment also may not come as a regular check. It might be a signup bonus, a gift card, a monthly bill credit, a discounted thermostat, free or cheaper EV charging, an annual performance payment, or additional “export credits” for energy sent back to the grid.

The most expensive devices, namely EVs and home batteries, are often what yield the greatest savings, which adds a barrier to entry for those who cannot afford these products in the first place. A smart thermostat program can be a low-stakes way to start.

You might have a variety of reasons for wanting to sign up, including supporting the overall health of the grid or avoiding the construction of a new power plant in your community. “There are not that many things that you can do that directly contribute to decarbonizing the electric supply, or to improving affordability, or to improving reliability, and this is just a clearly effective way to do that,” Frader-Thompson says. “And you get paid for it.”

In short, the best VPP program is not necessarily the one that pays the most. It’s the one that clearly tells you what it can control, how much money you’ll get, how easily you can say no—and how well it supports a community’s energy goals. Your home probably won’t feel like a power plant. But if your thermostat, car, or battery can bend a little when the grid needs it, your home can act like a small piece of one.

SUMMARYCanada is funding a major recruitment drive to bring 64 researchers to universities nationwide, including 48 from U.S. institutions such as Harvard, Yale, MIT, and the NIH. Backed by more than $364 million in government support, the initiative targets fields including AI, climate science, and medicine, with hires going to schools such as McGill, the University of Ottawa, and Western University.

Canada is recruiting 64 researchers to universities across the country (source paywalled; alternative source), including 48 from U.S. institutions such as Harvard, Yale, and MIT. The hires are backed by more than $364 million in government funding and are part of a broader effort to attract researchers in fields such as AI, climate science, and medicine. The New York Times reports: While scientists often shy away from political discourse, some are saying the Trump administration's assault on science is behind their departure. "I used to live in the country that I thought was the most enthusiastic about the prospects for science improving the human condition, of any country in the world," said Phillip Zamore, the chair of RNA Therapeutics Institute at the University of Massachusetts. "And I woke up one day and that wasn't true anymore." He has been recruited to McGill University in Montreal, which has also hired five other researchers, and will join the medical faculty. "If scientists don't stand up for truth, no one will," Dr. Zamore said.

The Canada-bound brain drain from U.S. institutions began last year as the Trump administration put forward policies that targeted foreign students, academic freedom and funding for equity-related programs. Kevin Hall, a nutrition scientist who left the U.S. National Institutes of Health last year, accusing federal officials of censoring his research on ultraprocessed foods, has been hired at the University of Ottawa. "While certain countries are cutting research and turning their back on academic freedom, we're doubling down on science," Melanie Joly, Canada's industry minister, told reporters at the announcement, in Vancouver, of the new university hiring. She billed it as the world's "largest talent attraction" project. The European Union has made a similar push. "Years from now, we will look back at today's announcement, and we will be able to seize the lasting impact of our choices," Ms. Joly said.

Unpredictable decisions about funding prompted Seth Guikema, a professor in civil and environmental engineering, who has specialized in natural hazards modeling at the University of Michigan, to look elsewhere. His work focuses on how climate hazards inequitably affect communities, and that work has become harder to fund, he said. "Every country sets its priorities in terms of what is going to get funded, and I think Canada has done a very good job of supporting research in areas that really matter to society," said Dr. Guikema, who will start at Western University in London, Ontario in January.

Google is imposing new memory-use limits on Android apps as the AI data center boom contributes to a broader memory chip shortage that could leave lower-cost phones with less RAM. Developers will have until February 2027 to meet new thresholds for memory and bitmap usage. To aid developers, Google is adding tools to flag apps that exceed the limits and help prevent slowdowns and crashes. TechCrunch reports: The company explains that the mobile industry is now facing "significant hardware supply constraints that are altering device memory availability," which can then, in turn, affect the consumer's experience with their devices. To address this, Google is now establishing new performance thresholds across several areas, like dynamic memory usage and bitmap usage. In addition, Google is adding code optimization requirements designed to prevent things like app slowdowns and crashes related to performance.

SUMMARYAnthropic introduced the Model Hardware Standard, a research preview designed to let AI agents control physical devices through a common interface and data format. The system aims to simplify scientific experiments by replacing custom integration software with standardized drivers, potentially cutting setup time from weeks or months to hours or minutes. Anthropic says the approach could help AI coordinate equipment such as robot arms, microscopes, cameras, and lasers across a network.

A researcher reacts in surprise as Claude figures out how to make a robot arm pick up an aluminum can.
Anthropic
arstechnica.com
A researcher reacts in surprise as Claude figures out how to make a robot arm pick up an aluminum can.

For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other data and actions that take place inside a computer. Anthropic is now aiming to change that somewhat with what it's calling the Model Hardware Standard (MHS), a set of standardized drivers designed to let AI agents easily interface with and control arbitrary devices.

For now, the "research preview" of the MHS effort is being sold mainly as a way to help scientists streamline the arduous process of creating the custom software integrations that are often needed to get disparate components of an experiment working in concert. MHS can provide a common interface and common format for data sharing between these devices, Anthropic says, allowing them to talk to each other across a network "without needing a bespoke 'translator' program in between." The standardized system could reduce weeks or months of exacting experimental setup down to "hours or minutes," Anthropic writes. An Anthropic graphic illustrating how MHS serves as a "translation" layer between AI agents and multiple types of devices. Credit: Anthropic

In a video posted alongside the announcement, Anthropic Technical Staffer Alek Kemeny says the MHS effort was inspired by observing neuroscientist Arco Bast work through an experiment on memory formation in the brain at the HHMI Janelia Research Campus in Ashburn, Virginia. Kemeny said Bast had worked out an interface to get the rotating laser beams, microscopes, cameras, and myriad other components of the experiment to coordinate through a common interface. "This idea could be used to have AI run any science experiment in the world," Kemeny recalls thinking at the time.

Read full article

SUMMARYOpenAI is testing a new Persistent mode for Codex that would keep the agent working until a user puts it to sleep and let it generate follow-up tasks across sessions. Code found in the product’s code base suggests the mode can use prior interactions and knowledge of the user to decide what to do next, while limiting changes outside the user’s system to cases approved by the user.

Wired reports that OpenAI is testing a new "Persistent mode" for Codex that would let the agent keep working until explicitly "put to sleep," proactively creating follow-up tasks for itself across sessions and using prior interactions and knowledge of the user to decide what to do next. Code alluding to the feature was spotted in the product's code base, though it has not been rolled out or officially announced yet. From the report: Persistent mode appears in Codex's "reasoning effort" menu, in which users can select the level of computing power, tokens, and time they want to allow for an AI model to "think" before answering a prompt. It seems to be one of OpenAI's most computationally intensive settings. When users have selected Persistent mode, OpenAI's code base reads that Codex will "continue working until put to sleep." That's a stark contrast to currently available modes, which will stop working on a task after a few minutes or hours, even if it's not complete.

In another file in the code base, OpenAI describes a feature within Persistent mode called "proactivity." This appears to be a type of system prompt for agents in Persistent mode, which are told that their work is not done when they finish answering a user's request. Instead, the agent is instructed to proactively create follow-up tasks for itself. The agent is capable of working on those tasks across sessions and using past user interactions and "knowledge of the user" to decide what to work on. It also has a tool to message the user without being asked but is told to send these sparingly.

The instructions also set limits for the agent, according to the file. The agent is told that Persistent mode does not expand what it is allowed to do and that altering anything outside the user's own system requires the user's approval first -- seemingly intended to limit how dangerous a persistent AI agent could be. The file sits in the shared core of Codex rather than in the code specific to the terminal, seeming to suggest the proactivity feature is intended for more than the command line tool.

Hugging Face is like an ever-evolving warehouse of open AI models.
Hugging Face / anucha sirivisansuwan via Getty Images
arstechnica.com
Hugging Face is like an ever-evolving warehouse of open AI models.

Nvidia is reportedly moving forward to acquire Hugging Face for $12.9 billion. The acquisition could help the hardware giant expand and fortify its deep integration with the wider AI industry.

Among other things, Hugging Face is a cloud repository for AI models, similar in some respects to what GitHub is for conventional computer software. Developers and researchers search for models that meet certain criteria, download and run them, and fine-tune them into different variants that then get uploaded back up to Hugging Face.

Read full article

SUMMARYAnthropic is expanding its support for scientists by offering 10,000 free or discounted Claude seats for one year through a new team plan for researchers. The company is also widening its AI for Science program beyond biology to support compute-intensive work in other fields, with up to $50,000 in credits per project. Researchers in approved academic or nonprofit labs can apply, and the company says it is working with the US government on a separate access program for life sciences professionals.

As Claude becomes increasingly capable at scientific research, we are focusing on building products and programs to support the research community. In June, we launched Claude Science, a product that integrates the tools that researchers most commonly use, produces auditable artifacts, and provides flexible access to computing resources. We have also continued to broaden our AI for Science program, which provides free credits to researchers working on high-impact scientific projects.

Starting today, we are announcing a significant expansion of these efforts. We are opening 10,000 seats for scientists around the world to access Claude subscriptions for free and at discounted rates for one year through our new Claude team plan for scientists. Standard seats will be free and premium seats with 5x usage limits will be available for $15 per month. Over the coming months, we intend to extend this program well beyond the initial 10,000 seats.

Alongside these subscriptions, we are also expanding the scope and scale of our AI for Science program. To date, we’ve largely supported scientists using Claude for biological sciences. We are now looking to offer credits for researchers in other scientific fields as well, including those working on ambitious, compute heavy-research, such of the kind that resulted in progress on the Riemann zeta function and Claude’s work on protein design.

By helping scientists access and increase their usage of Claude through subscriptions, credits, and products like Claude Science, we aim to radically accelerate scientific discovery.

Ways to access Claude

To register for our Claude team plan for scientists, please complete the verification form here. You must be a principal investigator or equivalent at an academic or nonprofit research institution to qualify; once verified, you can add the researchers in your lab to your plan.

As your lab makes use of your allotted credits and requires more usage than standard or premium plans provide, you can apply to our AI for Science program for up to $50,000 in credits per project. Any researcher is eligible to apply.

For now, researchers working in biology and chemistry will still be limited to our Opus-class models. Claude Fable models will continue to block professional biology and drug development queries because of their potential dual-use risks. We’re working in partnership with the US government to establish an access program for life sciences professionals to use Mythos-class models for life sciences research and development. We have now enrolled our first participants, and expect to share more and increase access soon.

Gemini Notebook’s Expert Intelligence feature
Image: Google

Google's AI note-taking app, Gemini Notebook, can now pull information from the books you've purchased. The new "Expert Intelligence" feature allows you to bring titles from Google Play Books directly into Gemini Notebook, which means you can ask questions about the material, as well as generate plans, infographics, AI podcasts, and more based on their contents.

During a briefing with The Verge, Google Labs editorial director Steven Johnson showed how you can use the tool to generate a recipe book using the information in Michael Pollan's Food Rules. In another example, Gemini Notebook applied the knowledge from Kim Scott's management-focus …

Read the full story at The Verge.

SUMMARYResearchers in the Department of Biology developed PottsMPNN, a machine-learning framework that generates protein sequences using physical principles of structure and stability rather than trying to mimic natural sequences. Published in PNAS, the model improves predictions of how mutations affect protein stability and helps design novel proteins that can fold into desired structures. The work aims to expand protein engineering for future biological applications.

When researchers use AI to design novel proteins, they hope to guide it to see that there are many potentially feasible options because different sequences of amino acids, the building blocks of proteins, could potentially adopt the same structure. Shown here are five sequence-structure pairs for a single protein, where the green structure is a natural protein, while the others were generated from designed sequences.
Image: Lillian Eden/Department of Biology, with sequence and structure elements courtesy of Foster Birnbaum.
news.mit.edu
When researchers use AI to design novel proteins, they hope to guide it to see that there are many potentially feasible options because different sequences of amino acids, the building blocks of proteins, could potentially adopt the same structure. Shown here are five sequence-structure pairs for a single protein, where the green structure is a natural protein, while the others were generated from designed sequences.

A protein’s function is determined by its structure, and structure — the way a protein folds — is determined by its sequence of amino acids, the building blocks of proteins.

Many methods for designing novel proteins, including examples that could bind to a disease-causing molecule in our cells, involve a two-step process: The structure comes first, and then a machine-learning framework generates a repertoire of sequences that could potentially adopt that structure.

In nature, many different amino acid sequences can fold into the same structure. At the same time, one amino acid sequence can potentially adopt different structures depending on the protein’s flexibility or a functional trigger. Therefore, when researchers use artificial intelligence to design new proteins, the challenge is to guide AI to “see” that there are many potentially useful answers — that many sequences can adopt the same fold

“For years, the field has measured success by asking whether a model can reproduce the protein sequence that evolution happened to select — our work shows that this isn’t the best metric for protein design,” says Amy E. Keating, Department of Biology head, Jay A. Stein (1968) Professor of Biology, professor of biological engineering, and senior author of a paper recently published in PNAS.

PottsMPNN, a new machine-learning framework developed in the Department of Biology, incorporates the physical principles that govern protein structure and stability, improving sequence generation and the ability to predict how mutations will affect a protein’s stability. In other words, the model has a better understanding of the sequence-energy landscape, meaning the relationship between the identity of each amino acid and the stability of the protein.

Adding this framework to a protein design pipeline will allow researchers to design structurally feasible proteins with sequences that don’t resemble those of any native protein.

“If we’re thinking about a completely novel, designed structure, there would be no native sequence to compare it to,” says graduate student and lead author Foster Birnbaum. “What we actually care about is how likely the generated sequences are to fold into the desired structures, how well the model understands the sequence-energy landscape, and how well it can predict the effect of mutations on the stability of the protein.”

Beyond the noise

In the same way that AI has recently powered some dramatic social changes, so too has machine learning impacted the pace and breadth of fundamental biological research. Only recently has it become possible to reliably use a computational model to generate a protein structure or sequence. Perhaps the most widely used model today, however, was released in 2022.

“For a field that’s moving as fast as machine learning in biology, that model has not been surpassed — we’ve been trying to understand why that is, and what it is about that model that makes it so useful,” Birnbaum says.

Birnbaum was first interested in strategic applications of something researchers call “noise,” or adding variations to a protein structure during training. Noise decreases the tendency of the model to overly mimic native sequences, increasing the diversity of structures for which it’s able to generate sequences.

PottsMPNN also uses a pairwise distribution to capture interactions between amino acids. The ability to account for the physical interactions between all 20 possible sequence options at a pair of positions in the protein is a key reason that PottsMPNN more accurately models the sequence-energy landscape than other methods.

Finally, Birnbaum says, they introduced sets of evolutionarily related sequences into training the PottsMPNN framework to teach the model how different sequences can adopt the same folded structure.

Birnbaum acknowledges that in trying to shift away from adhering to native sequences, incorporating evolutionary information is, in some ways, still a reliance on them. But PottsMPNN succeeded in demonstrating that as the model depends less and less on native sequences, structural compatibility and energy prediction, including for novel proteins, improve.

Protein design in the age of AI

“Once we can design any protein we want, that enables us to do a potentially scary amount of biological engineering,” Birnbaum says. “It’s a difficult task, but I’m really optimistic about this century’s progress in biology.”

Birnbaum hopes that the model could be further improved and fine-tuned for a specific task, which has in the past led to better predictions, for example, on the outcome or consequence of a particular mutation.

Ultimately, according to Keating, “Our methods move the field toward designing useful new-to-nature proteins for diverse applications while providing a stronger foundation for future advances.”

SUMMARYAnthropic introduced a research preview of the Model Hardware Standard, a shared specification that lets AI agents operate lab and manufacturing equipment through a common interface. Early partners including Genentech, the University of Washington, Carnegie Mellon, HHMI Janelia, QuEra, and Tetsuwan used it to automate tasks such as protein assays, qPCR monitoring, microscopy workflows, serial dilution experiments, and quantum laser recovery. The company plans to open source the standard after testing safety evaluations with research and industry collaborators.

We’re opening a research preview of the Model Hardware Standard (MHS), a shared specification for AI agents to safely operate physical devices, to a first group of scientific research labs and advanced manufacturers. MHS enables AI agents to operate multiple lab and manufacturing instruments, such as microscopes, liquid handlers, and robotic arms, in parallel, and perform intricate tasks ranging from routine drug discovery experiments to laser calibration on a quantum computer. The development of MHS began as a collaboration between Anthropic and HHMI Janelia Research Campus.

It typically takes a lab or manufacturing facility weeks, if not months, to set up and integrate their hardware. Most devices don’t communicate with each other, instead requiring specialists to build bespoke integrations. MHS reduces this integration work to hours or minutes. And by incorporating AI into these tools, MHS also helps researchers and engineers more readily orchestrate autonomous, round-the-clock experiments and workflows, with agents able to reason through each step in an experiment, update parameters in real time, and, in some cases, recover from hardware errors without intervention.

We’re sharing an early version of MHS with partners across science, robotics, electronics, and manufacturing so we can collaborate to build safety evaluations and develop best practices for AI systems operating physical equipment, ahead of making the standard open source. MHS works with any device that has a programmable interface. It is also model-agnostic, and any agent harness can access it using standard protocols, such as the Model Context Protocol. To apply for access to the research preview, head here.

How MHS works

Before and after the Model Hardware Standard (MHS).

Getting multiple devices in a lab or on a factory floor to communicate with one another can be challenging, even setting aside the added difficulty of integrating AI into the setup. Each device tends to have its own programming interface, and so far there has been no standardized way to integrate them. And once the devices are connected, there is no common way for them to share data with an AI agent, nor to let the agent operate them safely.

MHS addresses these challenges by introducing a standardized driver: software that translates between a computer’s operating system and a hardware device. The MHS driver uses a simple set of primitives—commands like “read” (for example, “get temperature”) or “write” (for example, “set temperature”)—that any hardware device can understand and act on. And it makes each device discoverable in a standard format, so that devices and agents can find each other and communicate across networks without needing a bespoke “translator” program in between.

The MHS driver also helps an AI agent understand how to use a device it has never seen before, giving it information about machine characteristics that may not be discernable from code alone (for example, the weight of a robot arm, which is important for knowing how to manipulate it safely). To date, much of this information has been stored in paper manuals, on a user’s computer, or as tacit knowledge. But the MHS driver contains tags that let the user write this information directly in natural language (users can either do this themselves, or by chatting to an agent that interviews them about their hardware setup). With the information from these tags, the MHS driver then automatically produces a reference file with information about a device’s general characteristics, such as what it can measure, what can be adjusted, and what safety limits will be enforced. This file gives the agent everything it needs to know to operate the device.

After the devices are connected and the agent knows how to use each one, the agent needs a way to control the hardware. For MHS, there are three such mechanisms: MCP, the command line interface, and code files (APIs). These work together to enable orchestration across multiple devices via a single line of code.

Once the agent can control the devices, it’s able to receive operating data from each one and supervise and direct the work at a high level. The agent can sequence steps across instruments, monitor results, and adjust parameters as conditions change in real time. When the agent needs to execute long-running tasks or operate devices faster than its online reasoning would allow, it can chain together driver commands from one or more devices in code files. This allows the devices to carry out operations themselves, without the agent needing to reason at every step.

As we’ve tested MHS, we’ve found that Claude interacts with experiments and hardware in an exploratory manner, much as a scientist would. For example, we observed Claude make an adjustment to a laser, observe the results through a camera to assess how its adjustment moved the laser beam, and repeat the process, seeking to understand the sequence of events. Claude then packaged what it learned into code files, writing a deterministic script that let it align the laser without having to reason at each step, so the whole process could run as a single command.

Early examples from MHS

We are only just beginning to see what people can do with frontier models and MHS, but our hope is that the standard can be of use to researchers, engineers, and other practitioners in speeding up the process of discovery and experimentation in any domain that uses devices with a programmable interface.

As we developed MHS, we shared it with a handful of labs and hardware manufacturers in biotech, robotics, quantum computing, and other fields. Across these early projects, we saw MHS reduce the time it took to integrate devices, make it possible to iterate faster in a variety of experimental settings, and assist with the live operation of machines and real-time fault detection. Below, our partners share the details of some of their early projects involving MHS.

Genentech: Implementing MHS for lab automation

Researchers at Genentech implemented and tested MHS as a proof of concept for automating the BCA protein assay, a standard procedure to measure total protein concentration in a sample, which requires coordinating across a liquid handler, a robotic arm, and a plate reader.

Read more

Genentech

Genentech: Implementing MHS for lab automation

Researchers at Genentech implemented and tested MHS as a proof of concept for automating the BCA protein assay, a standard procedure to measure total protein concentration in a sample, which requires coordinating across a liquid handler, a robotic arm, and a plate reader.

For half a century, Genentech has been tackling some of the most formidable challenges in science and medicine. In 1977, our scientists successfully produced somatostatin—a peptide hormone that regulates insulin and glucagon, growth hormone, and digestive tract functions in humans—in E. coli bacteria using recombinant DNA technology, proving that bacteria could be reprogrammed into bio-factories for medicines. Shortly thereafter, we synthesized recombinant human insulin, which in 1982 became the first genetically engineered therapeutic ever approved by the FDA. Since then, our commitment to basic research and patient care has pushed us to discover breakthrough therapies for cancer, multiple sclerosis, and other complex diseases.

This work requires rigorous experimentation in our drug discovery labs, often with large-scale automated systems capable of running high-throughput assays and testing many variables in parallel. These systems are made up of highly specialized lab robots—liquid handlers, robotic arms, microplate readers, and the like—that need to be carefully calibrated, iteratively tested, and supplied with complex programming logic in order to carry out experiments with precision. Currently, setting up these automated systems is a manual, time-consuming process that can take weeks or even months, limiting the number of scientific ideas our researchers can test.

To address this core bottleneck between experimental design and automated execution, we implemented MHS for lab automation with Anthropic. This open framework is designed to standardize AI-to-hardware communication, enabling scientists to interact with specialized lab robots using natural language and eliminating the need to write custom robotic code. Our ultimate goal is to build autonomous labs where AI handles the tedious, mechanical parts of experiment execution at scale so that our scientists can focus on the many creative aspects of accelerating drug discovery that rely on human judgement, such as experimental design, interpretation, decision making, and the invention of new lab approaches.

Automating the BCA assay as a proof-of-concept

We first implemented MHS on a large robotic workstation designed around a liquid handler. We wanted to see whether MHS could speed up the automation of the bicinchoninic acid (BCA) protein assay, a standard procedure used to measure total protein concentration in a sample. The procedure involves three instruments: a liquid handler to make precise fluid transfers, a robotic arm to move labware, and a microplate reader to measure optical absorbance, or how much light a sample absorbs at a given wavelength. We deployed MHS across all three devices, using Claude to orchestrate the protocol and act as a central communication hub for the hardware. All experiments were conducted in standard 96-well microplates, a staple of automated lab equipment.

The BCA assay involves handling liquids with different physical properties, ranging from simple aqueous reagents to viscous, foamy protein samples. In our setup, we used bovine serum albumin (BSA) at known concentrations as our protein sample to serve as a reliable standard. Because these fluids behave differently under pressure and flow, pipetting must be done extremely precisely to ensure that an exact volume of solution is transferred. For example, BSA solutions are viscous and form bubbles at high flow rates—the speed at which a liquid moves through a pipette tip—which directly impacts how accurately the solution is pipetted into the plate. Virtually all automated scientific experiments start with optimizing such fluid dynamics for each protocol.

The automated experiment workflow. First, a scientist describes the experiment to Claude in plain language. Then, Claude plans and orchestrates the run, drawing on reusable skills and a knowledge base. Every instruction passes through MHS, which serves as the standard interface for each device. MHS then operates each instrument (the liquid handler, robotic arm, and microplate reader) and streams its state back to Claude. The orange ring shows the part of the experiment Claude executed in a closed loop. It set a flow rate and transferred dyed liquid to a plate, sent the plate down the stack and read absorbance, then scored its own transfer against an expert’s and adjusted the flow rate, converging on water ≈ 140 µL/s (0.016) and viscous BSA ≈ 10 µL/s (0.181).

As a starting point, we gave Claude the standard BCA assay protocol to establish a baseline against which to assess improvements. In this first test, Claude executed the protocol steps, but it selected generic liquid handling parameters with the same flow rate for both aqueous and viscous solutions, which caused bubbles to form in the viscous solution, resulting in inaccurate liquid transfers. We then asked Claude to autonomously optimize fluid dynamics for both plain water and viscous protein samples (BSA). We prompted the model with an experimental design to optimize the liquid transfer flow rate, asking it to explore our expert-defined range of flow rates by conducting trial transfers with dyed liquid and taking absorbance readings with the microplate reader to determine the optimal flow rate for each liquid type. Claude also had access to a “ground truth” transfer, performed by an expert in the same plate, and we asked it to minimize the difference between the expert’s results and its own. After performing the transfers, Claude calculated the root mean square error (RMSE) to quantify how accurate it had been (the lower, the better, with zero being the perfect score; if it aimed for 100 microliters but dispensed 98, that 2-microliter miss would count against the score).

Claude independently executed these trial runs and analyzed the resulting plate reader data to get closer to the expert-performed transfers. For water, Claude concluded that a flow rate of ~140 µL/s was optimal (0.016 RMSE); for BSA, it arrived at 10 µL/s (0.181 RMSE)—parameters that our automation experts confirmed were reasonable for our setup. Ordinarily, performing this optimization requires an automation specialist to write custom programming logic for every single parameter set, iteratively analyzing the data until they find the right parameters.

Autonomous error recovery and the limits of current AI models

During the experiment, Claude encountered several unexpected errors, including tip pickup failures and fluid detection errors, but managed to recover on its own—a capability that current scientific instruments mostly lack. However, these experiments also highlighted the current limits of AI models. Although they excel at general-purpose reasoning, they still struggle with physical, chemical, and biological constraints, particularly when troubleshooting errors that call for real-world physical intuition.

An example of this type of limitation is the formation of bubbles during liquid handling. Although they may seem benign, bubbles create a cascade of challenges: if a protocol calls for aspirating 40μL of reagent but there are air bubbles in the liquid, the actual liquid volume transferred will be lower due to the space occupied by air. Furthermore, liquid-level sensors can trigger hardware errors when a pipette tip encounters foam instead of liquid; bubbles also distort the optical readings that are the final readout of the experiment.

Genentech scientists analyze plates for the presence of bubbles.

When it encountered runtime errors caused by bubbles during mixing, Claude’s default instinct was simply to retry the operation in the same plate well with different parameters. But this only agitated the fluid further and created more bubbles. Because Claude did not yet understand the underlying physics of the failure, we had to guide it towards parameters that handled the liquid more gently.

Once Claude was informed that the error code stemmed from physical bubbles in the liquid and that it needed to move to a clean well and reduce the number of mixing cycles in order to correct the error, it maintained that context for the rest of the run. We subsequently codified these takeaways into reusable liquid handling skills for Claude, which allowed it to select sensible default parameters for liquids with varying physical properties, reducing the number of liquid handling errors. These experiments highlighted the sorts of reasoning limits we can address by refining Claude’s software harness for lab automation.

Towards autonomous discovery

Although there is more work to be done to improve how Claude reasons about physical lab manipulations, this study proved to be a highly promising proof-of-concept. By assessing Claude’s decisions against our own domain expertise, we are generating the datasets we need to continuously improve models’ performance in automating lab experiments.

Going forward, we plan to evaluate Claude and MHS to orchestrate broader, end-to-end autonomous workflows in our drug discovery labs. We aim to build an autonomous discovery engine where scientists set the high-level biological intent, and AI agents help them coordinate the physical pipeline—generating hardware instructions, executing experiments autonomously, running closed-loop analysis, and delivering screen-ready models and screening data.

To expand our scope and impact, we’ll need to implement MHS on additional hardware, such as centrifuges, automated incubators, analytical instruments, and sensors. We’ll also need to tune the agent harness so it understands the nuances of working across drug discovery from molecules to live, sensitive cells. And we’ll have to integrate other, custom models that monitor and adaptively optimize experiments around the clock based on real-time data. With AI handling the routine tasks of maintenance, quality control, and environmental monitoring, our scientists can focus more on high-level experimental design, reasoning, and invention—moving us one step closer to accelerating the discovery of life-saving medicines.

Acknowledgements

We’d like to thank the Genentech scientists who contributed to this work, including Anupriya Tripathi, Matthew Bucci, Justin Nicola, and Corinne Gullekson.

University of Washington Baker and Pinglay labs: Bringing AI agents to the bench

Zihao Song, a PhD student in the University of Washington Baker and Pinglay labs, used MHS to build a dashboard to remotely monitor his instruments; an AI agent-supervised qPCR (which copies a target DNA sequence through repeated cycles of heating and cooling) that watches amplification curves and halts the procedure at the right moment; and an integration between a robotic arm and a liquid handler for collision-free plate handoffs.

Read more

University of Washington Baker and Pinglay labs

University of Washington Baker and Pinglay labs: Bringing AI agents to the bench

Zihao Song, a PhD student in the University of Washington Baker and Pinglay labs, used MHS to build a dashboard to remotely monitor his instruments; an AI agent-supervised qPCR (which copies a target DNA sequence through repeated cycles of heating and cooling) that watches amplification curves and halts the procedure at the right moment; and an integration between a robotic arm and a liquid handler for collision-free plate handoffs.

De novo protein design—building proteins that have never been seen in nature from scratch—has found a steadily increasing number of applications in medicine, environmental protection, and more over the past several years. Two things have held it back, however: cost and throughput. These days, designing a protein like PETase (the enzyme that breaks down plastic) can cost as little as $0.01. But testing that protein at the bench is slow and expensive, costing around $100 and requiring a week of labor per candidate—which adds up, given that we test 1,000 candidates at a time.

As a PhD student in the Baker and Pinglay labs at the University of Washington, I am working to develop high-throughput methods to reduce the cost per experiment and dramatically increase the number of de novo protein designs that we can screen at once. But working at that scale comes with costs of its own. Every round I run, whether a multiplexed design assay or an active learning campaign on enzyme activity, presents the same two challenges: monitoring status and capacity. Currently, our monitoring instruments sit in different corners of the lab, so I can’t easily see how a run is going without physically walking over to each one to check. When something fails partway through the run—the HPLC halts on an error, for example, or a liquid handler misfires and ruins a plate—I seldom discover it right when it happens. By the time I notice, hours may have passed, and the experimental sample is unusable. And I own only one of most instruments, so a single machine sets the pace for a whole round of testing, and I spend hours feeding it by hand. The PCR step is the worst culprit in this capacity crunch: it only handles one plate at a time, and I have to change the plates every 90 minutes (which is how I sometimes end up moving plates at 4 a.m. instead of sleeping).

The obvious fix is automation. But a research lab runs on flexibility, and that’s the one thing traditional automation cannot incorporate. A typical factory line might run one protocol 10,000 times, but my lab runs dozens of protocols a year, half of them new, which I must revise mid-run when the protein yield comes back far below what we assumed or a DNA assembly fails. Plus, my instruments come from different vendors, each with its own software, data format, and driver. Wiring them together is an integration problem that takes months to years and can cost anywhere from thousands to millions of dollars, putting it out of reach for most labs. There is no standard workflow to automate a protocol, and no affordable way to connect the instruments in most academic labs.

To explore a low-cost, low-effort route around both, I combined MHS with an AI agent and ran a few demos in my lab. MHS essentially gave the agent eyes, hands, and a sense of timing: it could see the status of every instrument, run each one, and coordinate them to work together.

Figure 1. Comparing an academic lab, an automated lab, and an MHS-based lab. Traditional labs run distributed instruments without a central scheduler. This is flexible but labor-intensive, with AI use limited to human-AI exchanges. Automated labs integrate instruments under a scheduler for near-autonomous operation, but they’re expensive and inflexible, keeping them out of reach of most academic labs, and AI-integrated versions are impractical beyond demos. MHS-based labs schedule all instruments through the standard, letting researchers monitor, control, and coordinate equipment; their AI-native architecture also lets agents actively participate in experiments.

Case study 1: Taking the lab remote

Figure 2. Monitoring instruments using MHS. Researchers can monitor the status of all connected instruments directly through the MHS dashboard or via an AI agent. (left) Screenshot of the MHS dashboard; (right) output from Claude Code after connecting to MHS.

Prior to MHS, I had to rove around the lab to monitor instruments. With MHS, instruments report their status to one dashboard, so I and my colleagues can check on the whole lab from a laptop, or even ask an AI agent from a mobile phone without setting foot inside (Figure 2).

This remote monitoring is especially helpful for experiments that demand sustained attention. Quantitative PCR (qPCR) is a good example. qPCR amplifies (i.e., copies) a target DNA sequence through repeated cycles of heating and cooling, with a fluorescent reporter that brightens as copies accumulate. DNA amplification follows an S-shaped curve: the copying doubles the target each cycle, so the signal stays flat while it is still faint, climbs steeply once there is enough to detect, then flattens again at the top of the curve as reagents run low and the copies stop doubling (the plateau). Letting the reaction run into that plateau distorts the DNA library, such that I can no longer glean accurate data about the final quantity of amplified DNA sequences. To avoid that, I need to watch the curve and halt the reaction at the right moment. This can take many hours and requires that I actively monitor the instrument’s screen.

MHS addresses this tedium, monitoring and analyzing the amplification curves as they come in and reporting back in real time. It identifies the curve pattern and, at precisely the right junctures, asks the researcher whether to stop or continue. When told to stop, it halts the reaction and advances the instrument to the next step: a 4 °C hold, which keeps the DNA from degrading so it stays usable for downstream work (Figure 3). With an AI agent and MHS watching the curve, we can now focus on setting up downstream sequencing reactions at the bench or analyzing library enrichment data from other experiments in the office.

Figure 3. Using an AI agent to monitor and control an experiment in real time via MHS. We worked with Claude Code to automate the execution of a qPCR protocol, transmitting the curve for each cycle to the chat box in real time for review. Upon receiving a stop command, the system halted the protocol and loaded a hold protocol. (All curves represent actual images from the interaction with Claude Code; some output has been truncated.)

Case study 2: Coordinating instruments through a plate handoff

Other experiments don’t need real-time monitoring, but they do require me to repeatedly load samples into a machine and take them out (for example, high-throughput DNA amplification, protein purification, and plate-based assays like ELISA). Loading a sample only takes a few seconds, but each run takes an hour or two, so I end up returning to the lab every hour just to swap plates.

In an effort to free ourselves from full days tethered to the bench, we used an open-source robotic arm built on LeRobot, instrumented with MHS, to safely coordinate sample loading across multiple instruments. As a demo, I reproduced one routine handoff for a high-throughput experiment run. In this process, a liquid handler dispenses reaction reagents into a plate; the robot arm then lifts the finished plate off the deck and moves a fresh one into place, and the liquid handler dispenses again into the new plate. Claude Code controls and coordinates both instruments through MHS, running each step only once the previous one finishes, so the two instruments never collide during the handoff.

The demo worked as intended. After the liquid handler finished dispensing, the AI agent picked up the completion signal and, about 10 seconds later, triggered the arm’s next move, lifting the plate off the deck. Across repeated tests, the two instruments never collided: the arm never moved before dispensing had finished, and the handler never started before the arm had cleared the plate. Meanwhile, I watched the whole run on my office computer without touching anything. Handing off this kind of coordination to an AI agent, within the safety standards built into MHS, points to a future where an agent chains many such steps overnight while the bench runs unattended.

Looking ahead

Setting MHS up was faster and easier than I expected, especially given how my earlier automation attempts had gone—weeks spent evaluating platforms, chasing vendor support, learning and building glue code between instruments, and finally giving up. Connecting six instruments through MHS took under a week, including the time I spent writing drivers for them. Once they were connected, the AI agent worked with the instruments without much fussing on my part: it discovered each device, read its status, and called its operations without my having to hand-hold the interface. For someone who has spent years working around instruments that don’t talk to each other, that changed my day-to-day more than I anticipated. The time I used to spend monitoring qPCR curves now goes to planning experiments, reading papers, and analyzing data, or sometimes just taking a nap and spending an hour in the sun.

These demonstrations are still just proofs-of-concept. More complicated experimental protocols will require significant optimization to work reliably, as well as the integration of broader and more complex physical manipulations. Running an agent continuously over long monitoring windows also has compute costs that need to be weighed against the researcher time saved.

These considerations aside, we are excited to continue to experiment with how MHS might help us run a fully autonomous design-build-test-learn round. Every round of de novo protein design or optimization currently stalls at the handoffs, where I carry results from one stage to the next; in the future, with MHS giving an agent a stable interface into every instrument, that cycle could run on its own. I can envision an agent proposing a set of designs, running the builds and assays, reading the results back through that same interface, and using what it learns to plan the next round. A lab that can generate its own scientific data in this way, round after round, is beginning to look reachable, even on an academic budget.

Acknowledgements

We thank peer reviewer Pushya Krishna as well as Xander Balwit, Rebecca Hiscott, Ethan Dyer, Conor Kelly, and Siddharth Mishra-Sharma for providing helpful feedback. We are grateful to Alek Kemeny and Bailey Bova for helping us set up MHS. Special thanks to Dr. Sudarshan Pinglay for his contributions, support, and guidance on the blog, and for Dr. David Baker’s mentorship in my research.

Carnegie Mellon University: Determining dose-response curves through rapid automation

Researchers at Carnegie Mellon University used MHS to run serial dilution dose-response experiments about three times faster than before, with an AI agent orchestrating a liquid handler, a plate reader, a robotic arm, and monitoring cameras spread across three computers with fundamentally incompatible interfaces.

Read more

Carnegie Mellon University

Carnegie Mellon University: Determining dose-response curves through rapid automation

Researchers at Carnegie Mellon University used MHS to run serial dilution dose-response experiments about three times faster than before, with an AI agent orchestrating a liquid handler, a plate reader, a robotic arm, and monitoring cameras spread across three computers with fundamentally incompatible interfaces.

Sina Barazandeh, Arth Banka, Gün Kaynar, Jiayi Li, Peneeta Wojcik, Carl Kingsford, Jose Lugo-Martinez, Joshua Kangas

A key component of drug development is determining dosage. Once we have identified a drug candidate, we need to understand how much of the drug is necessary to be effective—too much can be costly, or even toxic, and too little is ineffective. The appropriate dosage is usually determined through a process known as serial dilution. We start with a strong solution and dilute it by the same ratio each time, using the last dilution to make the next one. For example, mix one part solution with nine parts solvent to get a 10x-diluted sample; take one part of that sample and dilute it again, in the same way, and repeat. Each step lowers the concentration by a fixed amount, giving us an even, predictable range to test.

The process is time-consuming and error-prone, typically requiring multiple iterations to determine the right maximum concentration and the appropriate step size between dilutions. Too high a maximum concentration risks saturation, meaning that the signal maxes out and the curve flattens at the top, so those high doses stop providing any useful information about the response. Too small a step size doesn’t cover a wide enough range; too large a step size skips over the transition region entirely, missing the point at which the response actually changes.

When done by hand, setting up and conducting a set of serial dilution experiments can take weeks. Already onerous in traditional drug development, this is even more impractical in the high-throughput screening of AI-directed drug development, where the aim is to determine the dosages for numerous candidates at a time. It comes as little surprise, then, that serial dilution experiments are a prime target for robotic laboratory automation.

Unfortunately, setting up such automated experiments is itself complex and time-consuming. It requires coordinating multiple pieces of experimental equipment across several rounds of experimentation to obtain a usable dose-response curve. Even with access to an automated laboratory (a non-trivial requirement, given the need for multiple automation-compatible instruments and costly integration software), it can take weeks of automation engineering and protocol development to develop a procedure to carry out these experiments.

Our solution

MHS enabled us to run these experiments roughly three times faster by allowing AI to programmatically control several pieces of laboratory equipment. Our system combines a CyBio Felix liquid handler (a robot that moves precise volumes of liquid between wells, tubes, and plates), a Varioskan LUX plate reader (the instrument that measures an optical signal, such as fluorescence, in every well of a microplate), a robotic arm to move 96-well plates, and monitoring cameras with an AI-controlled orchestrator to automatically and dynamically measure dose-response curves.

Individually, each of the components is challenging to control programmatically, requiring a unique interface and specific operation modes. An engineer normally has to learn and hand-code a separate integration for every instrument before they can work together. Using MHS, however, we were able to develop drivers from scratch for each of these instruments and an orchestration layer that lets a Claude Opus 4.8 agent run the full protocol autonomously. This took about eight hours, versus the several weeks a vendor-built setup typically takes.

CMU laboratory instruments. (left) The Analytik Jena CyBio FeliX liquid handler for automated pipetting and liquid-transfer workflows. (right) The Thermo Scientific Varioskan LUX multimode plate reader for microplate-based absorbance, fluorescence, and luminescence measurements.
CMU laboratory instruments, continued. The Thermo Scientific Spinnaker robotic arm for automated microplate handling and transport (left), with monitoring cameras used to observe plate movement and system operation (center and right).
A 96-well plate arranged as a serial dilution, with the concentration decreasing step by step across the columns from 200 µg/mL to 0.20 µg/mL. This produces a broad, predictable concentration range that can be measured to build a dose-response curve.

Hardware, setup, and workflow

Our setup uses three computers. Computer 1 runs the robotic arm, controlled through scheduling software that takes job files dropped into a submission directory instead of a normal API. Computer 2 runs the liquid handler through an older Windows ActiveX/COM scripting interface, plus the monitoring cameras over USB. And computer 3 runs the plate reader, which has no programmatic interface at all, only an on-screen GUI. MHS turns each of these into one manifest of states (the conditions a system can be in; for example, plate at position 3, sample at 25°C, well filled) and procedures (the operations it can perform, such as aspirating or shaking), so the model works from a single, consistent interface, no matter which of the three computer control styles is running underneath.

The workflow itself is identical to what it was before MHS, only it’s now agent-driven: the liquid handler prepares a dilution series, a camera check confirms the plate is present and correctly oriented before any transfer is allowed, the arm moves the plate to the reader, the reader takes the measurement, and the model looks at the resulting curve and decides whether to adjust the concentration range and run it again or accept the result. To test it, we used a colorimetric dye (a dye whose color intensity tracks its concentration) as a stand-in for the actual drug candidate. This kept the experiment safe and easy to visualize while still requiring the same decision-making a real dose-response run would need.

Each instrument’s interface brings its own challenges. The arm’s scheduler is based on a directory watcher that generates two different files per submitted XML file, which MHS must reconcile to get one clean result, typically within a second of submission. The liquid handler only exposes COM scripting with no modern SDK, so each usable method had to be worked out either from vendor documentation or from a Claude Opus 4.8 agent exploring the interface to write a functional driver. A single dispense cycle takes about four to five minutes, and the allowed error margin on dispensed volume is only 5% before the resulting curve becomes unusable. The version of the plate reader software we use has no API of any kind, so MHS drives its GUI the same way a person would, with nothing to check its work against except what’s visible on screen.

Before this, a person had to sit through each of these steps: watching the arm’s log for failures, checking that the plate was seated correctly, and deciding whether a resulting curve was informative enough to keep or whether the concentration range needed adjusting and the whole thing needed to be rerun. MHS and the agent now handle all three of these decisions directly and automatically.

What we have achieved with MHS

To verify that MHS would operate safely and correct itself like a human operator would, we artificially induced six different conditions: missing plate, rotated plate, reader busy, disconnected camera, unreachable device, and active emergency stop. The system correctly blocked all six before any device moved. Then we asked the agent to run the serial dilution experiment to achieve an acceptable curve. The model evaluated the resulting curve, but found a fit too poor to accept (R² < 0.9, driven by saturation in the upper concentration range) and decided independently to discard the plate and rerun on a fresh plate with a compressed concentration range (200 µg/mL top concentration reduced to 100 µg/mL). The second run produced a strong, usable fit (R² > 0.98 with 3.4 variation across repeated measurements) with no human input at any point.

Run 1. The first serial dilution experiment tested concentrations up to 200 µg/mL. At the higher concentrations, the measurement began to saturate, meaning the signal stopped increasing in a useful way. Because this made the dose-response curve less reliable, the system rejected the run and decided that the concentration range needed to be adjusted.
Run 2. The system automatically repeated the experiment with a lower maximum concentration of 100 µg/mL. This new range captured the changing response much more clearly, producing a stronger and more reliable dose-response curve. The improved fit was accepted without any human intervention.

The most impressive thing about MHS was the integration speed. The time from raw, non-automated equipment readiness to a completed dilution curve, including one autonomous rerun, was eight hours. By contrast, engaging a vendor to deliver a working automated setup typically takes multiple weeks. The instruments run their own native software as usual; MHS adds an orchestration layer on top, with no additional automation software required. Any device with an API, SDK, or GUI interface can be integrated. The drivers developed for each instrument have been standardized and will be made publicly available, so others can reuse them rather than repeating the integration work from scratch.

What’s next

Future work in our lab will focus on validating the system with real drug candidates and replacing the dye’s color signal with readouts that capture actual biological effects. We also plan to expand MHS support to instruments such as qPCR and microscopes, and to integrate MCP-based agents with the MHS fleet. We also hope to reduce the integration time per instrument so we can scale the system for larger workflows.

We are also making sure this automation can be carried out safely. We plan to add more safety checks, monitor instrument and device responsiveness throughout our experiments, and refine our protocols for when and how human approvals are required for high-risk decisions. We’re looking forward to further exploring how else we can speed up the automation of our experiments, as MHS helps our researchers move faster, safely.

HHMI Janelia: Using MHS to accelerate microscopy research

At HHMI Janelia Research Campus, researchers are using MHS to speed up a range of microscopy-related projects. Here, Virginie Ruetten, a scientist in the Ahrens lab who studies how sleep helps the body recover from stress, shares how she used MHS to unify and orchestrate a rig that previously involved seven different vendor programs without a shared interface.

Read more

HHMI Janelia

HHMI Janelia: Using MHS to accelerate microscopy research

At HHMI Janelia Research Campus, researchers are using MHS to speed up a range of microscopy-related projects. Here, Virginie Ruetten, a scientist in the Ahrens lab who studies how sleep helps the body recover from stress, shares how she used MHS to unify and orchestrate a rig that previously involved seven different vendor programs without a shared interface.

A few nights of disrupted sleep are enough to cause widespread impairment: altered cognition, dysregulated metabolism, a weakened immune system. If this goes on long enough, sleep loss can even prove fatal. Yet we still don’t fully understand why. Part of the reason sleep is so hard to study is that it isn’t localized to any one organ. Because sleep is a whole-animal state, developing a mechanistic understanding of it requires measuring many parts of the body at once.

Microscopy offers a way to do this. Cells engineered to express fluorescent sensors emit light that signals their activity; microscopes can image these signals with high temporal and spatial resolution, letting us observe what these cells are doing. However, most animals are too large or too opaque for such imaging to function across the body. My work thus uses young zebrafish. This model organism is popular for its small size and transparency, and its organs and many aspects of its sleep physiology are similar to those of mammals, including humans.

These properties, combined with an experimental approach I developed called WHOLISTIC imaging, allow us to use two-photon microscopy to capture cellular activity throughout the brain and body of a living zebrafish. Rather than taking snapshots of isolated tissues, we can watch how cells and organs respond and interact from moment to moment across the entire animal, giving us a better understanding of the cellular players and underlying mechanisms behind physiological processes.

WHOLISTIC imaging of body-wide cellular activity in a larval zebrafish seven days post-fertilization. Maximum-intensity projection through the full volume of a young fish expressing the calcium indicator GCaMP7f in all cells, imaged with a customized mesoscope, a two-photon large field of view microscope. Fluorescence transients report intracellular calcium, a proxy for cellular activity, and are visible simultaneously in the brain, spinal cord, heart, gut, and peripheral tissue.

Controlling and coordinating devices through a unified interface

The instruments needed to carry out my experiments fill an entire room. As is common in many advanced microscopy setups, my rig is cobbled together from many components, each of which has been bought separately, from a different manufacturer, and wired up by hand: powerful femtosecond lasers; fast galvanometer mirrors, which sweep the microscope’s laser beam across the sample; super-sensitive photomultiplier detectors, which collect the returning light; and two precise translation stages, which position the fish relative to the sample holder, and the sample holder relative to the microscope.

These devices have to operate on a tight, shared schedule: the laser must be gated in step with the mirrors that scan it, and the stage must compensate if the animal moves or the sample drifts out of the focal plane. However, these devices were not designed to work together. Each comes with its own vendor control software, with no common interface. They often run in different programming languages, too: the detectors run in MATLAB, the cameras in Python, the electrophysiology in C#. A quantity held by one program, such as the position of the stage, is therefore unknown to the others. Yet the devices need to communicate—for example, each stage needs to know the other’s location for the system to know the absolute location of the sample.

The consequence of this incompatibility is that I spend a lot of time figuring out how to get devices to talk to each other, writing bespoke code to bridge two programs—or, in some cases, resorting to adding yet another device, a digital acquisition board (DAQ), a card that physically routes and transforms electrical signals, so the devices can communicate. Once everything is wired up, I still need to launch seven programs in a fixed order just to start an experiment, an error-prone process where getting the launch order wrong can cost the whole session.

MHS replaces those point-to-point connections with a single interface. Each device is now described and onboarded once, and its variables, controls, and sensor values are recorded in a single dictionary that lives in shared memory, a region of the computer’s memory that the operating system lets many programs access.

The benefit is that the cost of hardware integration stops scaling with the number of devices. Before MHS, integrating new pieces of hardware into the system was a multi-day project. Since implementing it, however, when I added a new camera to image the laser beam, it took me only a few minutes, and I could seamlessly feed the camera’s output—the location of the beam—back to the mirrors steering the beam, allowing me to more precisely align it. Starting an experiment now involves one click on the MHS dashboard instead of seven separate steps.

Beam alignment using MHS. Data from a laser beam camera streams through the MHS state dictionary. A digital target (white cross) can be added to guide alignment, and the beam can be precisely centered manually or using motorized mirrors controlled by an agent.

Quantitative monitoring and online analysis

Even after the hardware is wired up, I still need to run a plethora of checks and parameter adjustments before and during experiments to ensure I’m acquiring high-quality data. This includes monitoring the fish’s health, ensuring the camera is focused on the heart to measure heart rate variability, and surveying the quality of the fluorescence image to ensure that the cells I want to record are visible at high resolution. Each check involves computing a derived quantity from one of the many data streams the rig produces, such as the signal-to-noise ratio of the fluorescence data or the fish’s heart rate.

Such quantitative monitoring used to be laborious, as each data stream was collected by a separate program, and the values each program held in memory could not be easily read by any other program while the recording ran. Before MHS, I had three options, none of them optimal. First, I could collect the data and analyze it afterward, iterating on parameters between runs once it was saved to disc. But this took hours, and sometimes ended with the discovery that the recording was unusable. Second, I could judge the data by eye in the vendor viewer. This was fast, but it only gives an impression, not a precise measurement that can be compared across runs. Third, I could bolt analysis code onto the program doing the recording. But this was a pain because the code had to be rewritten for every program producing a data stream. Displaying the data streams was similarly time-consuming, because each application needed its own bespoke viewer, written in whatever language the recording program used.

MHS unified that fragmented process and removes the per-program rewrite. With MHS, each data stream is stored in shared memory in the MHS state dictionary, in a documented format that is readable by any process that attaches to it. Because each data stream is presented in the same way, analysis or visualization code can now be reused across devices and written in any language.

This allowed me to write a modular online analysis framework that guides the data through a chain of processing steps: data enters from a slot (an entry in the MHS state dictionary), passes through reusable transforms (operations on the data), and the result is written back to another slot or to disk. For visualization, I wrote a set of viewers—one per data type, rather than one per device—for images, time series, spectra, etc. Now, any data stream can be inspected while it’s being acquired, and I don’t need to rewrite any code to inspect a new one. When I became interested in how the zebrafish’s heart rate changes across the sleep-wake cycle, I could add a transform to compute the spectral content of the heart activity, as recorded by a camera, to estimate its heart rate; I could then reuse the code to compute the spectrum of the concurrent neural activity acquired by a different device, in a different language. Effort now compounds in one codebase rather than being split across one per device.

Online heartbeat tracking and prediction using MHS. Data from a camera imaging the ventral side of the animal is streamed through MHS, where it can be monitored using a generic MHS array slot viewer. Once in the MHS state dictionary, the data is instantly accessible to other processes, allowing for online identification of the heart and real-time tracking of heart activity (blue curve). Another process fits a predictive model enabling phase-locked stimulation (orange curve).

Running smarter experiments with agents

Experiments always involve tradeoffs. In imaging, for instance, I have to trade speed against coverage: I can scan a single plane—one thin optical slice through the brain—quickly, or many planes, to cover more of the brain, slowly. Finding the cells with the oscillatory activity I care about requires coverage, but measuring that activity requires speed.

Today, most experiments are ballistic: I select one set of settings at the start, one condition, and launch the run. So I have to pick a point on that tradeoff before the data tells me which point I need. Experiments last hours, so staying at the rig throughout is impractical—and biology is too variable and messy to have a simple deterministic algorithm do the searching for me. Ideally, I wouldn’t have to trade one for the other; I’d be able to search broadly, find the population of cells with the oscillatory activity I care about, then sample that precise region fast enough to resolve phase relationships between cells.

Agentic microscopy is the obvious way to get there: let an agent identify a region of interest and zoom in on it. But historically, that’s been easier said than done. The hard part isn’t getting the agent to iterate on writing analysis code to figure out where to zoom; it's getting it to reliably control a rig where commands move real devices, and where failure means a crashed objective, or an agent losing the few hours the sample preparation holds to the quirks of half a dozen vendor programs.

This is where I found MHS particularly helpful. With the entire rig’s state in a shared, standardized dictionary, agents can read and write every variable through a single interface, instead of seven vendor APIs, which removes the failure modes specific to each of them. I wrote a simple harness that made my rig operable by an AI agent. The core deterministic loop iterates between acquisition and analysis, and each result feeds the next decision. The agent enters at decision points, choosing the acquisition parameters (for example, what region to image) and what analysis to run, online and offline, in service of a user-stated goal. And because MHS enforces device-level safety limits, I don’t need to worry about the agent accidentally using excess laser power, for example, which risks bleaching the fluorescent molecules and degrading the sample. I'm still developing the framework and supervising experiments, but it has already allowed me to find the oscillatory population that a fixed setting would have missed, so I need fewer repeat runs and fewer animals to get the same number of usable recordings.

Going forward, I want to understand how these oscillations in the brain contribute to sleep and arousal so that my colleagues and I can identify targets for drugs that deliver restorative sleep, not just sedation. This requires mapping which cells are coupled to the oscillation, then running perturbation experiments phase-locked to it, to uncover the mechanism by which these cells shape global brain state. Doing this involves fitting models to the activity of thousands of cells as the data streams in, then triggering light-based activation of specific neurons—a technique called optogenetics—while the recording is running.

Such closed-loop experiments are hard because of the need to coordinate so many devices at speed, on a rig assembled from several vendors’ hardware that share no low-latency interface. MHS supplies the speed and adaptability, while still being easy for both humans and agents to comprehend. It doesn’t dispense with the tradeoffs that come with cellular activity imaging—such as speed versus coverage or how bright the signal is versus how long it lasts—since those are set by physics. But it does change how quickly I can explore the parameter space, home in on the right set, and iterate through experimental conditions and hypotheses. With MHS, I now look forward to the day when hardware control no longer limits the questions I can ask.

Online neural activity monitoring using MHS. Data from the two-photon microscope, imaging the hindbrain of a fish, is streamed to MHS, making it instantly accessible to other processes. The researcher can define regions of interest to monitor live activity (top trace: muscle; bottom trace: neuronal population), which is then fed back to an MHS slot, making it accessible for downstream processing.


Virginie Ruetten’s is just one of a handful of projects incorporating MHS at Janelia. Another team, led by Arco Bast in the Spruston lab, images neurons and their dendrites deep in the brains of living mice as they learn to navigate a virtual environment, watching memories form in real time. MHS grew out of Arco’s idea of putting the entire rig’s state in a standardized dictionary in shared memory, and his custom microscope was the first rig to run on it. Every laser, mirror, and sensor in his rig is exposed through MHS, so Claude can align the beams, tune the optics, and check its own results against the sensors, turning a half-day of manual setup into a single step. Another team, co-led by Magdalena Schneider and Hari Shroff, is using MHS to enable agentic control of a light-sheet microscope. This allows Claude to act as an orchestrator, deciding in real time how to image developing C. elegans embryos, and how to make trade-offs between competing imaging parameters. Across all projects, MHS has helped the researcher compress integration times, and made it possible for AI agents to control complex combinations of instruments, data visualization, and analysis, paving the way for faster discovery across a range of domains.

QuEra Computing: Using MHS in quantum laser stabilization

QuEra, a company that builds quantum computers using neutral atoms, used MHS to give an AI agent control over parts of the laser system inside its quantum machines. The agent developed a controller that recovers the laser’s “lock”—the ultra-precise frequency the lasers must hold to interact with the atoms—99.3% of the time without human intervention.

Read more

QuEra Computing

QuEra Computing: Using MHS in quantum laser stabilization

QuEra, a company that builds quantum computers using neutral atoms, used MHS to give an AI agent control over parts of the laser system inside its quantum machines. The agent developed a controller that recovers the laser’s “lock”—the ultra-precise frequency the lasers must hold to interact with the atoms—99.3% of the time without human intervention.

The promise of quantum computing lies in its ability to harness the properties of quantum mechanics to perform calculations that are out of reach of even the largest supercomputers. QuEra’s neutral-atom approach uses the naturally occurring quantum mechanics observable in a single atom as the foundation for these calculations. Nearly all aspects of the control, operation, and readout of these atomic qubits—the quantum bits that hold the computer’s information—are done through the controlled interaction of a laser with atoms.

For that to work, each laser has to hold its color (its frequency, measured in Hz) to an astonishing precision, roughly one part in a trillion. That’s equivalent to measuring the distance from the Earth to the Moon to within the width of a human hair. Physicists call a laser “locked” when its frequency is held this tightly. Everyday disturbances, such as temperature, vibration, or a shift in pressure, can push the laser off the desired frequency and “unlock” it, causing quantum operations to start to fail. As quantum computers perform long, error-corrected programs, a laser that ends up off its target frequency mid-run can spoil a computation that took hours to build up.

Figure 1. Part of the optical path that delivers laser light to the atoms.
Figure 2. The vacuum chamber where the atoms are held, inside the glass cell at center.

The laser at the center of this work is a titanium-sapphire laser, a tunable workhorse that atomic physics and quantum technology have relied on for more than 20 years. Traditionally, these lasers were controlled by hand, tuned for one-off experiments by experts. In QuEra’s quantum computers, those experts are aided by software that can detect and correct the most common disturbances before the lock is lost. But the potential sources of failure are so dynamic and varied that they cannot all be planned for in advance, so new or uncommon failures still need expert intervention. That requires an experienced operator, who must watch several instruments at once and judge what moved, what to correct, in which order, and when to trust the result. It typically takes 5 to 10 minutes to recover the frequency.

In university labs, this skill is handed down from one graduate student to the next, and when the lock drops at 2 am, someone has to wake up and drive in to perform the recovery. For a university lab, that is an inconvenience and an inefficiency. At the fleet scale of a quantum-computing company, it is simply untenable to keep this skill set in the hands of just a few people.

Figure 3. The laser system, and how Claude reaches it through MHS.

Given the vast and variable problem space, the complexity of the hardware to be controlled, and how critical this laser system is to the operation of the quantum computer, the team at QuEra identified automated laser recovery as a good first test case for MHS (Figure 3).

What we achieved with MHS

Making laser recovery automatic was not a new idea at QuEra. Before MHS, a team comprising a laser-systems engineer, a software engineer, an algorithms specialist, and a tester spent several months building a bespoke script to automate it. That recovery script reproduced what a human does at the bench, step for step: disarm the function holding the lock, work through the laser’s tuning controls in order, check the frequency after each tune, start over if the frequency is still off, and reengage the lock once it’s correct. But it only worked about 58% of the time, and it took around 150 seconds per attempt.

Because the script automates exactly what an expert does, it also inherits the same shortcomings. A linear sequence cannot absorb a change midway through, so when a shift in temperature or air pressure (for example, from somebody opening the door to the lab) undoes a step that had already succeeded, the recovery procedure must start over. This is the case both for an engineer at the bench and for the script, which is part of why a human recovery takes 5 to 10 minutes. Automating the steps made the sequence faster, but not fast enough to outrun the disturbances. This explains both the 150 seconds per attempt and the 42% failure rate.

QuEra handed the same problem to Claude through MHS. The team began by populating the agent’s context with a goal—write a standalone Python script to relock the laser—and a definition of success—relock on the first attempt, then hold for 30 seconds. The first time we attempted this, it took a day or two; now it only takes a few hours. We then induced disturbances for the agent to recover from: blocking the beam, cutting power to instruments to imitate a surge, and pushing the frequency off target by varying amounts. MHS supplied the access to Claude so it could read the instruments and move the controls as a human operator would.

Having set that up, the agent loop ran as four roles, each a fresh instance of Claude. One proposed a hypothesis for making recovery faster or more reliable; one wrote that change into the recovery script; one ran the updated script against the live laser and logged every step; and another read the logbook and decided what to change next (Figure 4). That cycle repeated hundreds of times, unattended, throughout the night, with each pass iteratively improving the script. By morning, recovery was taking about six seconds and working 96% of the time, against the 150 seconds and 58% it started from (Figure 5).

Figure 4. The loop the agent ran overnight.
Figure 5. Converging overnight. The 96% success rate shown here is from the development run; the 99.3% reported in the text is from the later blind test.

This improvement was the result of Claude rewriting that linear sequence as a decision tree. Instead of one path for every disturbance, the script it converged on reads each instrument, builds up if-then conditions from what the instruments show, and makes adjustments based on the specific physical disturbance and the pattern Claude learned from encountering it repeatedly. For example, if the frequency has barely moved, most of the laser’s controls will not change anything, so the script touches only one or two, leaving the rest alone. A human operator would still have to work through all of them, because the only way to be sure a control is right is to check it. Claude found the shortcut by running disturbances over and over again until the pattern in the lock’s behavior became clear. Claude’s advantage was its ability to do this exceedingly quickly, at a pace no operator could match.

We then tested the finished script against the same randomized set of induced disturbances with no agent involved. Across 700 trials, it recovered the correct lock 695 times, a 99.3% success rate. The hardest disturbances, where the frequency had hopped far from target, took 10 to 14 seconds, compared to the 5 to 10 minutes for a human at the bench, and the simpler ones took 0.9 to 5.4 seconds (Figure 5). The end product was a deterministic, fully inspectable script capable of running in production without an AI agent controlling it.

With the laser locking working significantly better, we then pointed the agent at the quality of the lock—that is, how often it unlocks. That’s set by 12 interdependent parameters (known as PID) inside the servo loop (shown in Figure 3). Tuning them well strips residual noise out of the lock, which sharpens the computer’s gates and makes unlocks rarer. Measuring that noise precisely requires capturing an oscilloscope trace and running a Fourier transform on it, which is not realistic for a human after every small change across 12 parameters. Instead, a specialist generally tunes against the root mean square (RMS) error the servo reports, which is a good enough approximation for lock quality. The PID parameters drift with temperature and pressure changes, so the goal is a tune that is good enough to hold for a while, followed by a retune when the lock starts slipping. The video below shows what just three of those parameters do.

Figure 6. What tuning entails for three of the 12 parameters. The video cycles through settings that overshoot, undershoot, and land.

The team pointed Claude at that same RMS number, but after every single adjustment it also captured a trace and computed the full spectrum hundreds of times over the course of the night. That is the part a human cannot match. Minimizing the RMS error usually does produce the lowest noise across the whole band, and a specialist who has done it for years is typically correct to trust it. But Claude did not have to trust it; instead, it could confirm the exact amount of residual noise after every change, and keep scouring through the search space until it could guarantee it had found the lowest realistic noise possible.

The laser’s parameters had already been set by a QuEra specialist, and that standing tune measured 15.7 mV of residual error. Over 363 experiments and 16 unattended hours, Claude brought it to 1.55 mV, roughly t10 times quieter on the RMS measure it was optimizing against (Figure 7). To verify that independently, the specialist retuned the same laser from scratch by his usual method, without seeing what Claude had found, and both sets of parameters went to a phase noise analyzer—a rare, specialized instrument used to measure absolute phase noise ( how much fluctuation is present at every frequency). This check also served as a way to fairly compare each PID parameter set. The agent’s tune matched the specialist’s across the band, with one exception: a roughly 220 kHz resonance where the manual tune had left about a thousand times more noise than Claude’s—exactly the kind of error that can occur when using the RMS heuristic, and what Claude’s method was able to avoid). The final check was the one that matters most in practice: how long each tune holds. Over a 19-hour run, Claude’s PIDs did not lose the lock once, while the expert tuned PIDs unlocked about 1.6 times an hour. Unlike the relock controller, this tuning workflow keeps the agent in the loop to adjust the parameters as conditions change.

Figure 7. The tuning run. Each dot represents a setting the agent tried out. The blue line shows the best overall result at each point in the experiment.

What’s next

The relock controller was built for a single laser system. But a quantum computer holds many laser systems with similar potential points of failure—not to mention many other subsystems that serve different functions but are similarly precise, fragile, and complex, and which require the same meticulous attention from a scant pool of experts.

In the pilot, MHS and AI agents did not replace such expertise completely. During the experiments, if something went wrong with the physical hardware, Claude didn’t know how to troubleshoot, as its understanding of the rig was programmatic rather than physical. Claude also often stopped to wait for human confirmation before performing an action it deemed even slightly risky, meaning experiments would sometimes pause overnight while Claude waited for approval. Still, an overly cautious agent is preferable to one that is not cautious enough. Finally, the team needed to provide a ton of context to Claude about what they wanted from the experiment and how Claude should carry it out for it to perform the tasks correctly.

Even with these limitations, the pilot showed that an AI agent can meaningfully improve how such systems are controlled, and some of these limitations should ease as models become more capable and build up a more sophisticated understanding of hardware. Next, QuEra aims to deploy the relock recovery system on live quantum processors, package the tuning workflow as a standalone tool, and apply the lessons from this experiment to other subsystems—working towards a fleet of machines that increasingly look after themselves.

Read more about the pilot experiment on the QuEra blog.

Tetsuwan Scientific: Using MHS to run qPCRs to profile local pollution

Researchers at Tetsuwan integrated MHS with its automated biology lab platform, ResearchOS. MHS helped orchestrate a qPCR workflow to contribute to citizen science efforts to characterize pollution in California’s San Pedro Creek.

Read more

Tetsuwan Scientific

Tetsuwan Scientific: Using MHS to run qPCRs to profile local pollution

Researchers at Tetsuwan integrated MHS with its automated biology lab platform, ResearchOS. MHS helped orchestrate a qPCR workflow to contribute to citizen science efforts to characterize pollution in California’s San Pedro Creek.

Since the 1960s, labs have had automated machines that can pipette, seal, shake, move labware, and perform most of the other functions of a biology lab. Yet the majority of biology experimentation remains manual. Part of the reason for this is that most biology experiments are fundamentally dynamic: sample count, plate format, and the number of conditions change between runs, as do the scientific parameters, such as dilution series depth, incubation times, the number of timepoints, and so on.

Translating a single experiment configuration into an automated workflow takes a specialist, known as an automation engineer, weeks or months. So automation only pays off when a configuration is repeated at enormous scale, like in high-throughput screening, where one method is used across a library of hundreds of thousands of compounds. The bulk of experimentation remains manual, inheriting all the potential errors and reproducibility problems that entails.

Tetsuwan is building an automated biology lab, available to researchers and agents via an API. Without a way to automate both small- and large-scale configurations, our lab would be confined to the limited capabilities of lab automation today. We built ResearchOS to solve this. ResearchOS is an automation platform that allows users to generate, run, and manage automated workflows without prior lab automation experience within minutes. Claude works with users to turn natural-language protocols into a script written in our syntax for experiments, which is ultimately processed by a custom compiler into automation code.

Built on top of a custom compiler, Tetsuwan’s ResearchOS makes it possible to execute automated experiments from natural-language instructions.

ResearchOS connects users to the automated lab, but the lab itself presents a formidable orchestration challenge. It is, after all, a menagerie of pipetting robots,1 robotic arms, and automated labware that all use different languages and all have their quirks. We implemented MHS to allow us to use Claude as an orchestration layer over that fleet by enabling these devices to communicate with one another and with the user. To test MHS, we put the platform to work on a citizen science project in Pacifica, California, running quantitative PCRs (qPCRs) to characterize sources of fecal contamination in the San Pedro Creek, which for decades have been at dangerously high levels.

Implementing MHS in a qPCR workflow

One of the most commonly used protocols in many biology labs, the polymerase chain reaction (PCR) copies DNA with nothing but a metal block, called a thermocycler, that heats up and cools down. DNA is a two-stranded molecule, and heating it separates the strands so they can be duplicated. Cooling the strands lets primers—short pieces of DNA that match the edges of the region you want copied—stick to those strands and mark where copying should start. Warming the strands back up, though not as much as before, lets an enzyme called a polymerase extend each primer along the strand it’s stuck to, copying it. qPCR adds a dye that glows brighter as copies accumulate, so we know how many copies there were to begin with. We can use this to measure how strongly a gene is switched on, to detect and quantify a pathogen in a patient sample, and, as in our case, to measure how much of a specific organism is present in an environmental sample.

A screenshot taken from the “procedure” page of ResearchOS, which shows users a graphical representation of their experiment after a protocol is uploaded.

qPCRs require a viscous, soap-like reagent known as a “master mix,” which contains the chemistry shared between reactions. These liquid properties make master mix especially prone to creating bubbles and foam when pipetted, which can cause inaccurate pipetting and degrade the quality of the experiment. With MHS, we were able to connect a camera that detects such errors and triggers an automated recovery process.

Claude, using MHS, operates a camera to take pictures of each transfer. The pictures are processed by a computer vision algorithm to identify pipetting errors, such as bubbles and foam. Claude can then intervene when an error is detected.

In one case, the camera identified bubbles in the master mix, which was in a tube held by a robotic arm. There was nothing the robotic arm alone could do to get rid of the bubbles. So ResearchOS scanned the lab for MHS-connected devices that could help, and Claude suggested an error handling strategy to us via Slack: move the tube to a centrifuge and briefly spin it at a low speed to draw the liquid to the bottom of the well, eliminating the bubbles. Through MHS, Claude was then able to issue the appropriate commands to the centrifuge.

This orchestration layer also allows protocols to stay hardware-independent. For example, a protocol might call for spinning a plate down at 15,000 × rpm for five minutes, without naming a specific type of centrifuge. ResearchOS can use MHS to query the network for a compatible centrifuge, learn its driver interface, and then use Claude to convert the force specified in the protocol into whatever parameters that specific machine accepts. For a machine that only takes rotor speed, for instance, that means dividing the force by the radius of the centrifuge’s rotor. The protocol author never even has to know which centrifuge was used, nor how the measurement was converted.

Multiple lab robots work in tandem to execute a qPCR workflow and troubleshoot errors.

Improving compiler heuristics with MHS

We also used Claude to set up a closed-loop optimization experiment to improve our compiler, which translates high-level code into instructions our lab equipment can execute. We took our qPCR protocol and compiled it into a range of realistic worklists that spanned a variety of experimental setups, such as testing different primer sets with different combinations of samples, sample numbers, and replicates (multiple runs of the same sample). From this, we determined how many different types of liquid transfers our compiler could implement. We then tried out those transfers on a robot under varying conditions, and measured their accuracy and precision with a tracer dye. When an experiment finished, Claude retrieved the accuracy data from the plate reader via MHS, analyzed and visualized it, and suggested tweaks to improve our compiler’s model of transfer precision. With more precise predictions and known systematic offsets (predictable errors in measured values), the compiler was able to make more informed machine layout and tip-reuse tradeoffs, which will ultimately help make our experiments faster, cheaper, and more accurate.

Over the course of this experiment we tested 9,143 individual dispenses, 300 unique transfer types (liquid × tip × volume × dispense count × over-aspiration), and 1,508 measured conditions across four types of liquid. On held-out experiments, the model Claude and MHS helped refine predicted multi-dispense precision roughly 12% more accurately than the manufacturer’s technical specification, beating it on 31 of 45 runs (sign-test p ≈ 0.001). This increased to roughly 17% on our most-replicated data.

Preliminary data from the San Pedro Creek

Although these were just early tests, the preliminary data from our qPCR testing corroborated the San Pedro Creek Watershed Coalition’s finding that humans are the primary contributor to fecal contamination in San Pedro Creek. We used qPCR to amplify fragments of bacterial DNA specific to different host organisms, such as humans, horses, birds, and dogs.

We were able to detect the presence of E. coli with a general 16S marker, as well as the presence of Bacteroides bacteria using the AllBac primer. 16S is a ribosomal RNA gene shared by all bacteria. By designing primer sets to interrogate sequences of 16S that contain interspecies variations, the host species can be identified (in this case, E. coli).

qPCR Amplification plot of HF183, a marker specific to bacterial species originating from human feces.
The qPCR experiment’s amplification curves show that the only source-specific marker detected was for humans (BacH, HF183). Note that Mean Cts are not comparable between markers.

We also saw clear amplification of human-associated Bacteroides using the HF183 and BacH primers. No other host-specific Bacteroides were detectable in our experiment. In future experiments, we aim to better characterize the source of contamination and quantitatively evaluate the concentration of human-associated Bacteroides in the wastewater.

Ultimately, the improved device integration, orchestration, and real-time error recovery made possible by MHS will be critical to bridging the gap between hardware, scientists, and models. As tools like ResearchOS and MHS improve the capabilities of lab automation, we believe experimentation will become increasingly accessible, reproducible, and programmable.

To read more about the optimization and community science aspects of this project, visit our blog.

1 Known as “liquid handlers” within lab automation. Note that 6-DoF arms are not the same as liquid handlers; liquid handlers are optimized specifically for pipetting rather than for general use. Workcells—which combine multiple lab robots into a single system—usually employ both liquid handlers (for pipetting) and robotic arms (for transferring plates and other labware).

01 / 06

Hardware vendors and the software companies that support them are also building MHS support into their equipment so AI agents can more easily discover and operate their devices. For example:

  • Amazon Web Services will support MHS through Strands Robots, the library for connecting AI agents to physical devices. AWS will provide participants a private, pre-release version of the Strands Robots package for the duration of the MHS research preview.
  • Automata is adding MHS support to LINQ, their lab automation platform, to perform intelligent error handling of instruments in autonomous labs.
  • Danaher and Anthropic are actively exploring how MHS-supported capabilities could enable its smart instruments and autonomous laboratories to scale biomedical research and development.
  • Doosan Robotics is testing MHS with their robotic arms, including to perform automated quality assurance and coordinate tasks across multiple robots.
  • MBF Bioscience is building an MHS driver for ScanImage, the software that runs laser-scanning microscopes in hundreds of neuroscience labs worldwide, to integrate AI agents into real-time data analysis and experiments.
  • QIAGEN is experimenting with MHS through a working proof of concept on its nucleic acid purification platform, QIAsymphony Connect, showing how AI agents could help laboratories troubleshoot instrument issues faster, guide operators through recovery, and improve instrument uptime while reducing risk to biological samples.
  • Tecan is adding MHS support for their Fluent liquid handling platforms so AI agents can discover and operate them directly.
  • Universal Robots has had early access to MHS and plans to add support to its robotics platform.

Joining the research preview

These early results from our partners are encouraging, but we have more work to do on the standard before we open-source it. As a large language model, Claude learns about the physical world through text and images, meaning its spatial and physical reasoning have limitations that still require expert oversight. When working with protein samples, for example, Genentech researchers had to guide Claude to recognize that errors caused by foaming in samples were physical failures, not software bugs, that could only be mitigated through the appropriate physical corrections.

MHS also doesn’t yet work with hardware that lacks a programming interface, so we’re working with the manufacturers of such devices to build in MHS drivers. Many developers already use Claude Code to work with individual pieces of physical equipment; for the next phase of MHS, we hope to expand the standard to cover more of the devices developers build on. Early adopters include Hugging Face, who are adding MHS support in LeRobot, their robotics library, and Raspberry Pi, who are enabling MHS integration across a number of their products following successful tests using their Camera MHS Driver.

We will also use the research preview to build additional safety evaluations with our launch partners and strengthen protections for the use of AI in the physical world. We are developing a physical safety roadmap to further bolster our safeguards policy and enforcement coverage against the risk of misuse. When we open-source MHS, we will release findings from the research preview as part of our guidance for deploying the standard safely.

We’re inviting stakeholders across industries to join the waitlist for our research preview of MHS. If you’d like to participate, submit your interest here.

Acknowledgments

MHS began as a collaboration between Alek Kemeny on Anthropic’s Beneficial Deployments team and Arco Bast, a postdoctoral scientist at HHMI Janelia Research Campus. Bast was running complex brain-imaging experiments on a rig that combined lasers, motorized focusers, and specialized cameras from different vendors with no common interface. To speed up his experiments, he developed a shared memory dictionary that enabled the instruments to communicate with one another at memory speed. Kemeny and Bast worked together to integrate AI models into that interface.

We thank everyone who has contributed to this work so far, including, but not limited to, Aaron Boswell, Ben Arthur, Boaz Mohar, Gagan Bhat, Mark Kittisopikul, Nadine Yasser, Nick Purcell, Takashi Kawase, and Virginie Ruetten. We look forward to moving MHS forward with our industry partners and, soon, with the open-source community.

SUMMARYNvidia has agreed to acquire Hugging Face for about $12.9 billion, according to reporting cited on Wednesday. If completed, the deal would place one of the most widely used open-source AI collaboration platforms under Nvidia’s ownership, strengthening the chipmaker’s position across AI software, models, and applications.

The Information reported on Wednesday that Nvidia has agreed to buy open-source platform Hugging Face for $12.9 billion. "Deal talks began after Hugging Face, an open-source AI platform developers use to collaborate, test and share tools, received acquisition interest from another suitor," reports CNBC, citing the (paywalled) report. Business Insider separately reported the acquisition talks. From CNBC: If completed, the acquisition would put one of the most widely used platforms for sharing and working with open-source AI models under Nvidia's ownership, expanding the chipmaker's reach further into the software and model ecosystem. Siddy Jobe, a fund manager at Eonopolis Exponential Technologies funds, said it made sense for Nvidia to target a company like Hugging Face, as Nvidia has made it clear that it is not looking to discriminate between closed and open-source models.

"I think Nvidia is very much a community, a platform-based company, and in that respect, I think Hugging Face fits perfectly within that. There is this five-layer cake from Nvidia, and foundational models are one of them," Jobe told CNBC's Squawk Box Europe on Thursday. "It is clear that Nvidia wants to be integrated in the entire stack vertically, going from energy to foundational models and also to applications," he added.

SUMMARYOne thousand Swiss citizens buried cotton underpants for two months as part of a citizen science project to measure soil health across Switzerland. The experiment, published in Plants People Planet, found that land management can strongly affect soil life and decomposition rates. The findings point to healthy soil as important for fertility, nutrient cycling, and ecosystem services.

After being buried in the ground for two months and broken down by soil organisms, a new pair of cotton underwear (left) shows clear signs of decomposition (right).
Nicolas Zonvi
arstechnica.com
After being buried in the ground for two months and broken down by soil organisms, a new pair of cotton underwear (left) shows clear signs of decomposition (right).

A thousand people voluntarily buried their cotton underpants for two months—not to keep the Underpants Gnomes from stealing them, but as part of a citizen science project to map out soil health in 25 countries around the world. The results of this unique experiment were reported in a new paper published in the journal Plants People Planet.

“Our results show that how soil is managed can significantly affect both soil life and soil quality,” said co-author Marcel van der Heijden, an agroecologist at the University of Zurich. “Healthy, biologically active soil is crucial for fertility, nutrient cycling and many other ecosystem services.”

Soil health is critical for agriculture and, by extension, food security, not to mention a healthy global ecosystem, but it has been declining worldwide, per the authors. One key indicator of soil health is decomposition rates of complex organic matter, breaking down stuff like leaves, wood, even cadavers into simpler organic and inorganic compounds, releasing C02 and nutrients in the process. There have been relatively few large-scale national decomposition studies involving different land use types, although smaller studies have shown that how humans use a site can significantly affect the soil's biological diversity.

Read full article

Three Microduck robots kicking a ball around
Image: Pollen Robotics / Hugging Face

Hugging Face's Pollen Robotics has launched its second cute AI robot, the Microduck, a one-eyed biped standing just under 10 inches tall. It's available to preorder now for $399 in cream, graphite, lavender, and sky blue, and Pollen Robotics says it plans to start shipping the little robot "before Christmas 2026."

Video demos of the Microduck show it picking up socks and markers, kicking around a ball, and zipping around on tiny rollerskates. Pollen Robotics says it can also "react to its surroundings" and follow around a laser pointer, or users can control it with a game controller. The Microduck's software is open-source, like Hugging Fa …

Read the full story at The Verge.

The Plaud One AI earbuds.
Image: Plaud

Plaud has introduced a new AI wearable that's designed to record, transcribe, and summarize your conversations, only this time it looks like earbuds instead of a pin. The Plaud One Explorer Edition can be worn like traditional earbuds or used through its standalone charging case, and the case includes built-in 4G to upload and process conversations without relying on your phone or Wi-Fi to stay connected.

Each earbud features 16MB of local storage (for a total of 32MB) and three microphones that can record at a distance of up to two meters. They can record for up to six hours according to Plaud, which is also the maximum estimated battery l …

Read the full story at The Verge.

SUMMARYNVIDIA announced major GeForce NOW updates at Gamescom 2026, including new DLSS 4.5 controls, broader device and browser support, and GOG single sign-on for easier cloud gaming access. The service is adding Steam Machine and Steam Controller support, expanding to more Amazon Fire TV devices, and bringing games like CONTROL Resonant, Gears of War: E-Day, Onimusha: Way of the Sword, The Blood of Dawnwalker, and STAR WARS Zero Company to the cloud at launch.

GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026blogs.nvidia.com

NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud.

New NVIDIA DLSS 4.5 technology controls give members more ways to fine-tune gameplay, while expanded support for new Steam devices, GOG single sign-on, Firefox browser support and more Amazon Fire TV devices continue broadening where and how GeForce RTX-powered PC gaming can be enjoyed.

Gamescom is also bringing a blockbuster lineup of PC games to GeForce NOW. CONTROL Resonant, Gears of War: E-Day, Onimusha: Way of the Sword, The Blood of Dawnwalker and STAR WARS Zero Company are all coming to the cloud at launch.

For a limited time, members can also get CONTROL Resonant with the purchase of a 12-month GeForce NOW Ultimate membership.

The Gamescom extras don’t stop there. To celebrate, GeForce NOW is giving away more than 40 prizes in our Community Giveaway — including gaming hardware and GeForce NOW Ultimate memberships.

There’s even more fun to come. Check out the list of 13 games joining the GeForce NOW library this week.

Next-Level Cloud Gaming

DLSS 4.5 games on GeForce NOW
More ways to choose how to play.

This fall, GeForce NOW is introducing new NVIDIA DLSS 4.5 controls in the cloud, giving Ultimate members more choice over how supported games look and perform, even without the latest hardware. Powered by NVIDIA RTX graphics, each control offers something different: DLSS Super Resolution sharpens image quality, while Dynamic Frame Generation delivers lower latency and stutter for smoother streaming, and Ray Reconstruction enhances ray-traced visuals. Together, these technologies make it easier to tune supported games for image quality, responsiveness or a preferred balance between the two.

Stay tuned to GFN Thursdays for details about best practices for enabling DLSS 4.5 overrides and how Dynamic Frame Generation can help Ultimate members enable more graphic settings while still getting smooth streaming.

GeForce NOW and Steam hardware
Stream full steam ahead with the Steam ecosystem.

GeForce NOW is also adding community-requested support for Steam Machine with Steam Controller later this fall, and later for Windows and macOS, delivering more ways to play and making it even easier for longtime Steam players to jump into the cloud using familiar hardware. The additions build on existing Steam Deck support and expand compatibility across the Steam ecosystem.

GOG SSO on GeForce NOW
Get your GOG on with one click into the library.

Launching GOG games just got even easier. With GOG single sign-on now available on GeForce NOW, members can securely link their account once, automatically sync supported games and launch them without having to sign in to GOG again — all with a single click.

With fewer steps between opening the app, skipping any downloads and jumping into gameplay, it’s another way GeForce NOW continues streamlining access to PC gaming libraries, making it easier to play the GOG library of games on TV platforms since gamers are automatically logged into their GOG accounts.

Find popular titles, like Cyberpunk 2077 and The Witcher 3: The Wild Hunt, in the app in the dedicated GOG row. Or search and filter for GOG games, then jump into supported titles from almost anywhere — all streamed from the cloud. Check back on GFN Thursdays for the latest games joining the GeForce NOW library, including newly supported games streaming from GOG.

GeForce NOW and Firefox
Outfox nearly any device with GeForce RTX gaming.

Plus, GeForce NOW is now officially supported on Firefox, expanding browser access across all major Windows browsers. Members can jump into supported PC games directly from Firefox without downloading a dedicated app. Ultimate members can enjoy GeForce RTX-powered gaming at up to 1440p and 120 frames per second.

Simply download the latest version of Firefox and visit play.geforcenow.com to get started. The cloud handles downloads, installs and updates, keeping local storage free while games stay ready to play in the cloud.

GeForce NOW and Fire TV
Bring the fire to living rooms with GeForce NOW.

GeForce NOW is expanding support to additional Amazon Fire TV devices. GeForce NOW first launched on Fire TV devices earlier this year, and will soon expand to the Fire TV Stick 4K Select, a 4K streaming stick that’s ready to game right out of the box. Streaming from the cloud means supported PC games are ready to play using a compatible controller, making it easy to enjoy GeForce NOW on the biggest screen in the house with even more options to play from the comfort of the couch.

And that’s not all the awesomeness from Amazon. Later this year, Fire TV users will be able to purchase GeForce NOW memberships directly through Amazon, creating a more seamless path from device setup to gameplay.

These new features will arrive alongside some of the biggest games coming to the cloud throughout the year.

Get Ready for Launch

Five newly announced and highly anticipated games are coming to GeForce NOW at launch: CONTROL Resonant, Gears of War: E-Day, Onimusha: Way of the Sword, The Blood of Dawnwalker and STAR WARS Zero Company.

Better yet, members can lock in one of them early: For a limited time, purchase a 12-month GeForce NOW Ultimate membership and receive CONTROL Resonant at launch at no additional cost, and be first to experience Faden’s next chapter. The offer is available Tuesday, Aug. 25, through Sunday, Sept. 27.

Ultimate members can stream the full lineup with GeForce RTX 5080-class performance in the cloud, including NVIDIA DLSS 4, ray tracing, NVIDIA Reflex and Cinematic Quality Streaming technologies, with up to 5K high dynamic range on supported devices.

Whether diving into new worlds or returning to iconic franchises, GeForce NOW lets members skip expensive hardware upgrades, massive installs and storage management so they can jump straight into the action the moment each game launches. Here’s a closer look at what’s coming:

Control Resonant bundle on GeForce NOW
Contain the chaos.

Explore a warped Manhattan on the brink of paranatural annihilation in Remedy’s CONTROL Resonant. Harness Dylan Faden’s extraordinary powers to battle the Hiss, the Mold and other reality-bending threats while searching for his sister, Federal Bureau of CONTROL Director Jesse Faden.

Gears of Wars E-day launch to be on GeForce NOW
Emergence begins.

Experience the horror and brutality of Emergence Day in the Coalition’s Gears of War: E-Day. Fourteen years before the original Gears of War, join Marcus Fenix and Dominic Santiago in a gripping origin-story campaign as the Locust Horde first erupt from below, igniting a desperate fight for survival. Rally squads and fight on in multiplayer with a reimagination of Gears’ iconic PvE mode, Horde Siege, or go head to head with a refined versus PvP.

Onimusha Way of the Sword launch on GeForce NOW
Pick up the Oni Gauntlet.

Fight across a twisted vision of Edo-period Kyoto in Capcom’s Onimusha: Way of the Sword. Wield the mystical Oni Gauntlet, master intense swordplay and battle relentless demonic Genma beneath mysterious clouds of Malice.

Blood of the Dawnwalker launch to be on GeForce NOW
Walk the line between day and night.

Step into 14th-century Europe in Bandai Namco’s The Blood of Dawnwalker, where war, plague and rising vampire powers have reshaped history. Play as Coen, a young Dawnwalker torn between humanity and cursed strength, and shape his path while fighting to save his family.

STAR WARS Zero Company launch to be on GeForce NOW
Command the galaxy’s finest.

Lead an elite squad behind enemy lines in STAR WARS Zero Company, a single-player turn-based game set during the twilight of the Clone Wars. Command former Republic officer Hawks and Zero Company through tactical battles and high-stakes missions where strategy, squad composition and player choices shape the fate of the galaxy.

More Ways to Win

Gaming gear and Ultimate memberships are up for grabs.

GeForce NOW is marking Gamescom with a community giveaway featuring more than 40 prizes. Gamers will have a chance to score a MacBook Neo, Steam Deck or Chromebook — each paired with a one-year GeForce NOW Ultimate membership — along with Amazon Fire TV Sticks and GeForce NOW Ultimate memberships ranging from one month to one year.

Head over to the GeForce NOW community on Reddit for all the giveaway details and how to enter.

The Hunt Is On

Squad up and fight for survival in Aliens: Fireteam Elite 2, streaming on GeForce NOW this week at launch. Assemble a four-player fireteam of Colonial Marines and battle through swarms of Xenomorphs, deadly new threats and increasingly desperate encounters. With deeper squad mechanics, new and improved classes, smarter enemies and expanded customization, every mission puts teamwork — and nerves — to the test.

Members can also look for the following games joining the cloud this week:

  • Aliens: Fireteam Elite 2 (New release on Steam, Aug. 25)
  • Resonance: A Plague Tale Legacy (New release on Steam and Xbox, Aug. 27, available on Game Pass)
  • STAR WARS Zero Company (New release on EA app and Steam, Aug. 27)
  • Baldur’s Gate 3 (GOG)
  • Call of Duty: Modern Warfare 4 – Open Beta Weekend (Battle.net, Steam and Xbox available from Aug. 28 10am PDT to Sep. 1 at 10am PDT. For more information please visit the Call of Duty website).
  • Clair Obscur: Expedition 33 (GOG)
  • Cold Fear (GOG)
  • IRON NEST: Heavy Turret Simulator (Steam)
  • Kingdom Come: Deliverance II (GOG)
  • Mount & Blade: Bannerlord II (GOG)
  • No Man’s Sky (GOG)
  • Paradise Killer (Steam)
  • The Witcher 3: Wild Hunt (GOG)

Dates listed above reflect when games are released on their respective stores. GeForce NOW availability may vary, as games are onboarded after release and added throughout the week. Keep an eye on GeForce NOW channels and GFN Thursdays for availability updates on announced titles.

To experience GeForce NOW, gamers can start with a GeForce NOW day pass and try premium cloud gaming before committing to a membership. Better yet, the cost of a day pass can be applied toward a first membership purchase, making it even easier to level up when ready.

Check back every GFN Thursday for new games, new features and even more ways to play on GeForce NOW. Which Gamescom announcement has you most excited? Let us know on X or in the comments below.

The future of PC gaming is on display with GeForce NOW.

🖥DLSS 4.5 controls
📱Expanded device support
Upcoming AAA games
🎮 More ways to play
🎁 Community giveaway
🦊 Firefox browser support
… and more! #gamescom

Details https://t.co/Rz6RnX0Pt4 pic.twitter.com/sShKRhzrD3

- 🌩 NVIDIA GeForce NOW (@NVIDIAGFN) August 25, 2026

The AI Assisted Editor in Photoshop
Image: Adobe

Adobe is rolling out an AI-heavy update for Photoshop that includes a new "optional" interface dedicated to its AI tools. Launching in beta, the "AI Assisted Editor" view will show all of Photoshop's AI features in a single toolbar, including its prompt-based image editor, background remover, an AI image extender, and more.

There are also new ways to refine edits with AI, including a "markup" feature to draw directly on an image to show Photoshop's AI assistant what you'd like to change. That means you can "select areas to recolor, sketch arrows to indicate position, or brush in rough shapes to suggest new elements" without using a text pro …

Read the full story at The Verge.

SUMMARYSlate Auto has unveiled a compact electric pickup called the Slate, designed to sell for under $25,000 with a 65-kWh battery and an estimated 205-mile range. The truck strips away many common features, with options such as power windows and Bluetooth sold separately, and deliveries are slated for late 2026 from a factory expected to build 100,000 vehicles in its first year. Investors including Jeff Bezos have put nearly $1.4 billion into the company, while Ford is also developing a smaller EV truck aimed at a 2027 debut.

EVs account for under 10% of total new-vehicle sales in the US, and the numbers are declining. From a climate perspective, that’s pretty dismal, especially because the transportation sector is the single biggest source of greenhouse-gas emissions in the country.

One thing that could help turn that around? Slate Auto’s new truck—a vehicle that seems to buck every convention about selling cars in the fully loaded, range-obsessed US market.

The company is going all in on simplicity, to the point of austerity. The Slate is a tiny, two-door pickup that’s shorter than a Honda Civic. Much of the media coverage has obsessed over the fact that the base model’s windows use hand cranks, a feature straight out of the 20th century.

Slate is also breaking away from other EV manufacturers’ efforts to compete with gas-powered vehicles on range. The truck sports a small lithium iron phosphate battery, ringing in at 65 kilowatt-hours, and its quoted top range is just 205 miles. For comparison, the most basic Tesla Model 3 can go over 320 miles on a charge.

By accepting a shorter range and forgoing the frills that Americans have come to expect in vehicles, Slate is able to offer the base model of its truck for less than $25,000. Some customers will choose to upgrade their truck with optional add-ons, like a Bluetooth stereo system, vinyl wraps, or even power windows. But the price will still likely come in well below the roughly $50,000 average for a new vehicle in the US.

It might seem an odd choice to go so small and simple in a market that’s increasingly sizing up—but the status quo hasn’t exactly been working for EV makers. A few years back, the hero for US automotive electrification was supposed to be the Ford F-150 Lightning. Announced in 2021, it was an electric version of the country’s best-selling vehicle.

But Ford discontinued the truck in December 2025, just four years after its introduction. It’s not entirely the Lightning’s fault: The second Trump administration slashed tax credits and other support designed to boost EVs.

The Lightning was also plagued by price increases. The base model cost roughly $40,000 when shipments started in 2022; during its final year, prices topped $54,000. One factor behind the increase was its massive battery; to reach an almost 300-mile range, the truck needed a battery with a capacity roughly twice that of Slate’s truck.

That obsession with range is largely unwarranted. While Americans have historically chased distance, the average driver puts under 35 miles a day on the odometer, and nearly 90% of trips in a personal vehicle are 20 miles or less.

Surveys of EV drivers show a similar trend: One recent study found that people tend to use less than 20% of their EV’s range on a typical day.

The notion that drivers need far less range than they think they do might sound a lot more persuasive these days—even for Americans who tend to buy cars for their longest road trip instead of their everyday errands. Roughly half the country is struggling to pay for basic necessities such as gas and groceries, and 95% of Americans believe we’re in an affordability crisis, according to a recent Harris poll. An inexpensive vehicle that allows you to skip the gas station sounds like an attractive prospect.

Affordable EVs have already found willing buyers in other parts of the world—even with reduced range. China currently has over 40 million EVs and plug-in hybrids on the roads, and roughly half of new vehicles sold are electric. On average, new EVs sold in China in 2025 had a range of just 247 miles. For the US over the same period, the average was 329 miles. (Europe falls between the two, at 281 miles.)

Slate is set to start delivering on preorders in late 2026. Its factory will have a capacity of 100,000 vehicles in the first year and 150,000 soon after. Thousands of customers have already put in preorders, according to the company.

Some people obviously believe in the truck’s prospects: Investors, including Jeff Bezos, have put nearly $1.4 billion into the company over three major funding rounds. Ford is jumping (back) into this space soon too. The automaker is working on its own small electric truck, which is expected to debut in 2027 at a retail price of around $30,000.

The key question that will determine Slate’s success or failure is whether drivers can get on board with a small, short-range vehicle for the sake of its price tag. At this moment, I’d bet the answer is yes.

Nvidia has reportedly agreed to buy Hugging Face, the popular open source AI hub, for $12.9 billion in a move that would let Nvidia both protect its chip empire and jump back into the cloud business.

An image of Nvidia’s logo
Image: Cath Virginia / The Verge

Nvidia's predicting it will pull in $108 billion in revenue within just a few months. It wouldn't be the first company to rake in over $100 billion in quarterly revenue - Amazon, Apple, and Alphabet have repeatedly reached the milestone.

Nvidia said in its latest earnings report that it brought in a record $96.2 billion in overall revenue in the past quarter, a jump of over $10 billion from the previous quarter. Its data center revenue alone more than doubled year-over-year to a record $89 billion, and the company's profits more than doubled to $59.7 billion.

A screenshot of a bar graph showing Nvidia's overall revenue in Q2 2027

Image: Nvidia

A screenshot of a bar graph showing Nvidia's Q2 2027 data center revenue

Image: Nvidia

Nvidia's "edge computing" category, which includes its consumer gam …

Read the full story at The Verge.

SUMMARYNVIDIA expanded NVLink Fusion with NVHBM, a custom high-bandwidth memory technology designed to improve performance and efficiency for semi-custom AI infrastructure. The system is said to deliver up to 30% higher memory bandwidth, 15% lower HBM power use, and up to 25% more compute-die area, while being offered through multiple memory partners. Amazon’s Annapurna Labs will be the first partner to work on the technology as part of its NVLink Fusion collaboration with NVIDIA, starting with future Trainium chips.

NVIDIA NVLink Fusion Expands With NVHBM Custom High-Bandwidth Memoryblogs.nvidia.com

The next wave of AI is placing new demands on infrastructure.

As AI agents and trillion-parameter workloads become mainstream, the performance of AI infrastructure depends not only on compute, but on how compute, memory, storage, networking and software are designed together as a unified system.

To help hyperscalers and AI innovators build the next generation of semi-custom AI infrastructure, NVIDIA today expanded NVIDIA NVLink Fusion with NVIDIA NVHBM, a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs. It will be validated and offered by leading memory partners, extending this advanced memory capability to NVLink Fusion customers.

Traditional HBM architectures place the memory controller on the XPU die, consuming valuable silicon area that could otherwise be dedicated to compute. NVHBM, built on the same technology that NVIDIA will use for future GPUs, integrates NVIDIA’s custom memory controller into the HBM base die.

By integrating the memory controller into the 3D HBM stack instead of the XPU, NVHBM delivers up to 30% greater memory bandwidth and 15% lower HBM power consumption, and frees up to 25% more area on XPU compute die compared with standard HBM4E.

NVIDIA is establishing a standard NVHBM implementation, available from multiple memory providers. This reduces the engineering effort required to integrate and qualify memory across multiple suppliers — giving NVLink Fusion customers a faster path for bringing custom AI chips to market.

Amazon’s Annapurna Labs will be the first to work on NVHBM as part of its broader collaboration with NVIDIA around NVLink Fusion.

AWS and NVIDIA Continue NVLink Fusion Collaboration

Amazon’s Annapurna Labs will work with NVIDIA on NVHBM technology and the NVLink scale-up architecture to enhance performance and efficiency for AI workloads.

This builds on AWS’s previously announced support for NVLink Fusion. Annapurna Labs will support NVLink Fusion with its next-generation Trainium chips starting with Trainium4, which would allow Amazon chips and NVIDIA GPUs to work together with common rack-scale architecture.

“NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency,” said Nafea Bshara, vice president of Annapurna Labs at Amazon. “We look forward to this technology collaboration to benefit future AWS infrastructure designs.”

Vertically Integrated and Horizontally Open

NVLink Fusion enables partners to connect custom XPUs and CPUs to NVIDIA’s rack-scale platform.

Partners can access NVIDIA NVLink chiplets, NVLink-C2C, NVLink Switches and NVIDIA MGX systems and racks, as well as a broad ecosystem of CPU partners, ASIC designers, system manufacturers and technology providers.

Offered with each generation of NVIDIA’s rack-scale system architecture, NVLink Fusion allows hyperscalers and AI-native companies to focus engineering resources on XPU innovation while using a proven technology stack for scale-up and scale-out networking, rack-scale systems and software — creating a faster, lower-risk path to deploying semi-custom AI infrastructure.

Learn more about NVLink and NVLink Fusion.

Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text
Aurich Lawson
arstechnica.com

While we wait (possibly in vain) for Gemini 3.5 Pro to launch, Google is releasing a different model in the 3.5 branch. The company has announced Gemini 3.5 Transcribe, an AI model designed to streamline voice input by editing out "ums" and corrections, outputting polished AI text. This model already powers the Gboard "Rambler" feature on the Pixel 11, but it's about to appear throughout the Google ecosystem.

According to Google, Gemini 3.5 Transcribe is much faster and more accurate than its previous voice-to-text engine, known as Chirp 3. The new AI model should be about 70 percent faster from voice to final transcribed text, and the live-speech error rate has dropped to 5.5 percent. That's only a little better than Chirp 3, which Google measures at 7.32 percent. Still, it's a pain to fix typos when you're using voice input, so any improvement here is beneficial.

Credit: Google

Read full article

SUMMARYMicrosoft announced a disc-to-digital program for Xbox owners that will let many Xbox One and Xbox Series X disc-based games unlock a digital entitlement by inserting the disc once. The feature, starting testing for Xbox Insiders on August 31, will support digital play, Xbox Play Anywhere, and Xbox Cloud Gaming for eligible titles while keeping the physical disc usable. Not all games will qualify at launch, and the entitlement can be revoked if the disc changes hands.

An anonymous reader quotes a report from Ars Technica: For decades, console owners have faced a choice between the convenience of digital downloads and the permanence of physical game discs. Soon, Xbox owners will be able to get the best of both worlds for thousands of supported titles as part of a newly announced disc-to-digital program. The program -- announced today ahead of testing for Xbox Insiders starting August 31 -- will let players claim a "digital entitlement" for "most Xbox One and Xbox Series X disc-based games" simply by inserting the disc into a console and launching it. That game will then be playable completely digitally, without the need to ever insert the disc, as long as you (or a member of your family account) is logged in. The digital entitlement will also allow access to features like Xbox Play Anywhere (for play on PC) and Xbox Cloud Gaming, for supported titles.

Microsoft says that your physical disc will "continue to work exactly as it always has" after the digital entitlement is claimed. But before you get any ideas, the fine print on the announcement mentions that there is only "one revokable license per game disc," so if you resell that disc or loan it to a friend, that digital entitlement could be transferred to a new account when someone else puts it into their console.

A leaked memo obtained by the Verge earlier this month suggests that publishers have to actively opt in to allow their game discs to activate digital entitlements, which could explain why "most" but not all Xbox One and Series X titles are supported. Windows Central separately suggests, based on discussion with unnamed sources, that "some discs may be incompatible due to how they were manufactured at the time," which could explain why original Xbox and Xbox 360 discs are not being discussed for the program. "While not every title will be available at launch, this is an important step toward a future where players can have greater confidence that the games they buy remain with them for years to come," said Xbox Vice President Jason Ronald.

SUMMARYAmazon has agreed to acquire DuckLabs, the company behind the open-source DuckDB database, and bring its team into Amazon Web Services. DuckDB will remain free under the MIT license and overseen by the independent DuckDB Foundation, while co-founders Hannes Muhleisen and Mark Raasveldt continue leading development from Amsterdam. Amazon plans to use the team’s analytics expertise to expand how customers analyze data in S3.

Amazon has agreed to acquire DuckLabs, bringing the team behind the popular open-source DuckDB database into AWS. "The deal fits Amazon's broader push to turn S3, its flagship cloud storage service, into a place where customers analyze data rather than just store it," reports GeekWire. "It gives Amazon a team experienced in building fast, lightweight analytics software that runs directly against data sitting in cloud storage." The DuckDB project itself will remain free and open source under the MIT license, overseen by the independent DuckDB Foundation. From the report: Employees of DuckLabs will join Amazon Web Services, including co-founders and DuckDB creators Hannes Muhleisen and Mark Raasveldt, who will continue leading the team and setting the project's technical direction. They will remain based in Amsterdam, where the team will continue developing DuckDB and related projects. [...] Financial terms were not disclosed. Amazon said it has signed a definitive agreement and expects the acquisition to close shortly. DuckLabs said it expects to become part of AWS in early September. Jordan Tigani, the CEO of MotherDuck, which sells a cloud service built on DuckDB, sees Amazon's acquisition as a predictable move to turn DuckDB's growing popularity into an AWS business. "That's Amazon's playbook, after all: wait until an open source project gets big enough, then launch it as a service," he wrote in a blog post, adding that "they're not acquiring Duck Labs just because they love open source."

Tigani also believes the deal could ultimately strengthen DuckDB, since Amazon has an incentive to keep the project open and widely adopted: "If DuckDB becomes the standard, it is going to drive a lot more compute on their infrastructure, which is where they make their money."

Vector illustration of the Google Gemini logo
Image: Cath Virginia / The Verge

Google has updated Gemini Audio with new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Transcribe is a new addition to the Gemini family that follows the launch of 3.5 Live Translate, and comes as we're still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out in June.

Google says that 3.5 Transcribe "represents a major advancement from our previous transcription model, Chirp 3," especially regarding multilingual performance and wording error rates. The transcription model allows users to "edit naturally with just your voice," according to Goog …

Read the full story at The Verge.

Mozilla plans to enable JPEG XL decoding by default in Firefox 157, which is due at the end of September. Phoronix reports: Firefox Nightly has JPEG-XL support enabled by default right now to help in vetting this support while Mozilla believes the support is in good enough shape for a stable debut with Firefox 157. This follows Chrome shipping JPEG-XL and Google Research developing jxl-rs as a Rust-based JPEG-XL image decoder that is both performant and secure.

In today's Mozilla Hacks blog post they elaborate on JPEG-XL vs. AVIF image formats: "JPEG XL: Excels at lossless imagery, progressive rendering, and further compressing JPEGs without quality loss. AVIF: Excels at web-quality photographic images, and images that have a mix of sharp edges and flat surfaces. ... Although AVIF tends to produce smaller files at web-quality than JPEG XL, AVIF only has basic progressive rendering support. So, for very large images, it may be worth taking the filesize hit with JPEG XL."

Loading more cached articles...