Feed aggregator

Covariant Announces a Universal AI Platform for Robots

When IEEE Spectrum first wrote about Covariant in 2020, it was a new-ish robotics startup looking to apply robotics to warehouse picking at scale through the magic of a single end-to-end neural network. At the time, Covariant was focused on this picking use case, because it represents an application that could provide immediate value—warehouse companies pay Covariant for its robots to pick items in their warehouses. But for Covariant, the exciting part was that picking items in warehouses has, over the last four years, yielded a massive amount of real-world manipulation data—and you can probably guess where this is going.

Today, Covariant is announcing RFM-1, which the company describes as a robotics foundation model that gives robots the “human-like ability to reason.” That’s from the press release, and while I wouldn’t necessarily read too much into “human-like” or “reason,” what Covariant has going on here is pretty cool.

“Foundation model” means that RFM-1 can be trained on more data to do more things—at the moment, it’s all about warehouse manipulation because that’s what it’s been trained on, but its capabilities can be expanded by feeding it more data. “Our existing system is already good enough to do very fast, very variable pick and place,” says Covariant co-founder Pieter Abbeel. “But we’re now taking it quite a bit further. Any task, any embodiment—that’s the long-term vision. Robotics foundation models powering billions of robots across the world.” From the sound of things, Covariant’s business of deploying a large fleet of warehouse automation robots was the fastest way for them to collect the tens of millions of trajectories (how a robot moves during a task) that they needed to train the 8 billion parameter RFM-1 model.

Covariant

“The only way you can do what we’re doing is by having robots deployed in the world collecting a ton of data,” says Abbeel. “Which is what allows us to train a robotics foundation model that’s uniquely capable.”

There have been other attempts at this sort of thing: The RTX project is one recent example. But while RT-X depends on research labs sharing what data they have to create a dataset that’s large enough to be useful, Covariant is doing it alone, thanks to its fleet of warehouse robots. “RT-X is about a million trajectories of data,” Abbeel says, “but we’re able to surpass it because we’re getting a million trajectories every few weeks.”

“By building a valuable picking robot that’s deployed across 15 countries with dozens of customers, we essentially have a data collection machine.” —Pieter Abbeel, Covariant

You can think of the current execution of RFM-1 as a prediction engine for suction-based object manipulation in warehouse environments. The model incorporates still images, video, joint angles, force reading, suction cup strength—everything involved in the kind of robotic manipulation that Covariant does. All of these things are interconnected within RFM-1, which means that you can put any of those things into one end of RFM-1, and out of the other end of the model will come a prediction. That prediction can be in the form of an image, a video, or a series of commands for a robot.

What’s important to understand about all of this is that RFM-1 isn’t restricted to picking only things it’s seen before, or only working on robots it has direct experience with. This is what’s nice about foundation models—they can generalize within the domain of their training data, and it’s how Covariant has been able to scale their business as successfully as they have, by not having to retrain for every new picking robot or every new item. What’s counter-intuitive about these large models is that they’re actually better at dealing with new situations than models that are trained specifically for those situations.

For example, let’s say you want to train a model to drive a car on a highway. The question, Abbeel says, is whether it would be worth your time to train on other kinds of driving anyway. The answer is yes, because highway driving is sometimes not highway driving. There will be accidents or rush hour traffic that will require you to drive differently. If you’ve also trained on driving on city streets, you’re effectively training on highway edge cases, which will come in handy at some point and improve performance overall. With RFM-1, it’s the same idea: Training on lots of different kinds of manipulation—different robots, different objects, and so on—means that any single kind of manipulation will be that much more capable.

In the context of generalization, Covariant talks about RFM-1’s ability to “understand” its environment. This can be a tricky word with AI, but what’s relevant is to ground the meaning of “understand” in what RFM-1 is capable of. For example, you don’t need to understand physics to be able to catch a baseball, you just need to have a lot of experience catching baseballs, and that’s where RFM-1 is at. You could also reason out how to catch a baseball with no experience but an understanding of physics, and RFM-1 is not doing this, which is why I hesitate to use the word “understand” in this context.

But this brings us to another interesting capability of RFM-1: it operates as a very effective, if constrained, simulation tool. As a prediction engine that outputs video, you can ask it to generate what the next couple seconds of an action sequence will look like, and it’ll give you a result that’s both realistic and accurate, being grounded in all of its data. The key here is that RFM-1 can effectively simulate objects that are challenging to simulate traditionally, like floppy things.

Covariant’s Abbeel explains that the “world model” that RFM-1 bases its predictions on is effectively a learned physics engine. “Building physics engines turns out to be a very daunting task to really cover every possible thing that can happen in the world,” Abbeel says. “Once you get complicated scenarios, it becomes very inaccurate, very quickly, because people have to make all kinds of approximations to make the physics engine run on a computer. We’re just doing the large-scale data version of this with a world model, and it’s showing really good results.”

Abbeel gives an example of asking a robot to simulate (or predict) what would happen if a cylinder is placed vertically on a conveyor belt. The prediction accurately shows the cylinder falling over and rolling when the belt starts to move—not because the cylinder is being simulated, but because RFM-1 has seen a lot of things being placed on a lot of conveyor belts.

“Five years from now, it’s not unlikely that what we are building here will be the only type of simulator anyone will ever use.” —Pieter Abbeel, Covariant

This only works if there’s the right kind of data for RFM-1 to train on, so unlike most simulation environments, it can’t currently generalize to completely new objects or situations. But Abbeel believes that with enough data, useful world simulation will be possible. “Five years from now, it’s not unlikely that what we are building here will be the only type of simulator anyone will ever use. It’s a more capable simulator than one built from the ground up with collision checking and finite elements and all that stuff. All those things are so hard to build into your physics engine in any kind of way, not to mention the renderer to make things look like they look in the real world—in some sense, we’re taking a shortcut.”

RFM-1 also incorporates language data to be able to communicate more effectively with humans. Covariant

For Covariant to expand the capabilities of RFM-1 towards that long-term vision of foundation models powering “billions of robots across the world,” the next step is to feed it more data from a wider variety of robots doing a wider variety of tasks. “We’ve built essentially a data ingestion engine,” Abbeel says. “If you’re willing to give us data of a different type, we’ll ingest that too.”

“We have a lot of confidence that this kind of model could power all kinds of robots—maybe with more data for the types of robots and types of situations it could be used in.” —Pieter Abbeel, Covariant

One way or another, that path is going to involve a heck of a lot of data, and it’s going to be data that Covariant is not currently collecting with its own fleet of warehouse manipulation robots. So if you’re, say, a humanoid robotics company, what’s your incentive to share all the data you’ve been collecting with Covariant? “The pitch is that we’ll help them get to the real world,” Covariant co-founder Peter Chen says. “I don’t think there are really that many companies that have AI to make their robots truly autonomous in a production environment. If they want AI that’s robust and powerful and can actually help them enter the real world, we are really their best bet.”

Covariant’s core argument here is that while it’s certainly possible for every robotics company to train up their own models individually, the performance—for anybody trying to do manipulation, at least—would be not nearly as good as using a model that incorporates all of the manipulation data that Covariant already has within RFM-1. “It has always been our long term plan to be a robotics foundation model company,” says Chen. “There was just not sufficient data and compute and algorithms to get to this point—but building a universal AI platform for robots, that’s what Covariant has been about from the very beginning.”

Diminutive Deep Sea Drone Dives for Wrecks and Reefs

The global ocean is difficult to explore—the common refrain is that we know less about the deep ocean than we do about the surface of the moon. Australian company Advanced Navigation wants to change that with a pint-sized autonomous underwater vehicle (AUV) that it hopes will become the maritime equivalent of a consumer drone. And the new AUV is already getting to work mapping and monitoring Australia’s coral reefs and diving for shipwrecks.

The Sydney-based company has been developing underwater navigation technology for more than a decade. In 2022, Advanced Navigation unveiled its first in-house AUV, called Hydrus. At less than half a meter long, the vehicle is considerably smaller than most alternatives. Even so, it’s fully autonomous and carries a 4k-resolution camera capable of 60 frames per second that can both capture high-definition video and construct detailed 3D photogrammetry models.

Advanced Navigation says Hydrus—with a depth rating of 3,000 meters, a range of 9 kilometers, and a battery that lasts up to three hours—is capable of a wide variety of missions. The company recently sold two units to the Australian Institute of Marine Science (AIMS), the country’s tropical marine science agency, which will use them to survey coral reefs in the North West Shelf region off the country’s west coast. Hydrus has also recently collaborated with the Western Australian Museum to produce a detailed 3D model of a shipwreck off the coast near Perth.

“If people can go and throw one of these off the boat, just like they can throw a drone up in the air, that will obviously benefit the exploration of the sea.” —Ross Anderson, Western Australian Museum

After many years of supplying components to other robotics companies, Peter Baker, subsea product manager at Advanced Navigation, says they company spotted a gap in the market. “We wanted to take the user experience that someone would have with an aerial drone and bring that underwater,” he says. “It’s very expensive to get images and data of the seabed. So by being able to miniaturize this system, and have it drastically simplified from the user’s point of view, it makes data a lot more accessible to people.”

But building a compact and low-cost AUV is not simple. The deep ocean is not a friendly place for electronics, says Baker, due to a combination of high pressure and corrosive seawater. The traditional way of dealing with this is to stick all the critical components in a sealed titanium tube that can maintain ambient pressure and keep moisture out. However, this requires you to add buoyancy to compensate for the extra weight, says Baker, which increases the bulk of the vehicle. That means bigger motors and bigger batteries. “The whole thing spirals up and up until you’ve got something the size of a minibus,” he says.

Advanced Navigation got around the spiral by designing bespoke pressure-tolerant electronics. They built all of their circuit boards from scratch, carefully selecting components that had been tested to destruction in a hydrostatic pressure chamber. These were then encapsulated in a water-proof composite shell, and to further reduce the risk of water ingress the drone operates completely wirelessly. Batteries are recharged using inductive charging and data transfer is either over Wi-Fi when above water or via an optical modem when below the surface.

Hydrus AUVs are charged using induction to keep corrosive seawater from leaking in through charging ports.Advanced Navigation

This has allowed the company to significantly miniaturize the system, says Baker, which has a drastic impact on the overall cost of operations. “You don’t need a crane or a winch or anything like that to recover the vehicle, you can pick it up with a fishing net,” he says. “You can get away with using a much smaller boat, and the rule of thumb in the industry is if you double the size of your boat, you quadruple the cost.”

Just as important, though, is the vehicle’s ease of use. Most underwater robotics systems still operate with a tether, says Baker, but Hydrus carries all the hardware required to support autonomous navigation on board. The company’s “bread and butter” is inertial navigation technology, which uses accelerometers and gyroscopes to track the vehicle from a known starting point. But the drone also features a sonar system that allows it to stay a set distance from the seabed and also judge its speed by measuring the Doppler shift on echoes as they bounce back.

This means that users can simply program in a set of way points on a map, toss the vehicle overboard and leave it to its own devices, says Baker. The Hydrus does have a low-bandwidth acoustic communication channel that allows the operator to send basic commands like “stop” or “come home,” he says, but Hydrus is designed to be a set-and-forget AUV. “That really lowers the thresholds of what a user needs to be able to operate it,” he says. “If you can fly a DJI drone you could fly a Hydrus.”

The company estimates for a typical seabed investigation in water shallow enough for human divers, the Hydrus could be 75 percent cheaper than alternatives. And the savings would go up significantly at greater depths. What’s more, says Baker, the drone’s precise navigation means it can produce much more consistent and repeatable data.

To demonstrate its capabilities, Hydrus’ designers went hunting for shipwrecks in the Rottnest ships graveyard just off the coast near Perth, in Western Australia. The site was a designated spot for scuttling aging ships, says Ross Anderson, curator at Western Australian Museum, but has yet to be fully explored due to the depth of many of the wrecks.

The Advanced Navigation team used the Hydrus to create a detailed 3D model of a sunken “coal hulk”—one of a category of old iron sailing ships that were later converted to floating coal warehouses for steamships. The Western Australian Museum has been unable to identify the vessel so far, but Anderson says these kind of models can be hugely beneficial for carrying out maritime archaeology research, as well as educating people about what’s below the waves.

Advanced Navigation used its new Hydrus drone to create a detailed 3D image of an unidentified “coal hulk” ship in the Rottnest ships graveyard off the western coast of Australia.

Advanced Navigation

Any technology that can simplify the process is greatly welcomed, Anderson adds. “If people can go and throw one of these off the boat, just like they can throw a drone up in the air, that will obviously benefit the exploration of the sea,” he says.

Ease of use was also a big driver behind AIMS’s purchase of two Hydrus drones, says technology development program lead Melanie Olsen, who is also an IEEE senior member. Most of the technology available for marine science is still research-grade and a long way from a polished, professional product.

“When you’re an operational agency like AIMS, you typically don’t have the luxury of spending a lot of time on the back of the boat getting equipment ready,” says Olsen. “You need something that users can turn on and go and it’s just working, as time is of the essence.”

Another benefit of the Hydrus for AIMS is that the drone can operate at greater depths than divers and in conditions that would be dangerous for humans. “Its enabling our researchers to see further down in the water and also operate in more dangerous situations such as at night, or in the presence of threats such as crocodiles or sharks, places where we just wouldn’t be able to collect that data,” says Olsen.

The agency will initially use the drones to survey reefs on Australia’s North West Shelf, including Scott Reef and Ashmore Reef. The goal is to collect regular data data on coral health to monitor the state of the reefs, investigate how they’re being effected by climate change, and hopefully get early warning of emerging problems. But Olsen says they expect that the Hydrus will become standard part of their ocean monitoring toolkit going forward.

This story was updated on 11 March 2024 to correct the year when Advanced Navigation unveiled Hydrus.