It was a lunch run that could have been completely unremarkable. A handful of hungry engineers, a craving for In-N-Out Burger, and a car that needed to get from the parking lot to the drive-thru window. But this particular trip had a twist that made it feel less like a quick bite and more like a small glimpse of the future. Aditya Ramabadran, Simon Mahns, and Tobias Gessler, three AI engineers working for a startup called Axiom, found themselves sitting in a 2024 Toyota Corolla near a Bay Area fast-food restaurant. They had no intention of driving themselves. Instead, with a laptop open and a conversation interface running, they asked OpenAI’s GPT-6 Astra to take the wheel. The car was equipped with several windscreen-mounted cameras and linked to its power steering system. A safety driver sat ready, one foot hovering over the brake, just in case the whole experiment went sideways. The AI model, which normally spends its life generating text, code, and the occasional image, slowly and steadily guided the vehicle up to the take-out window. They collected their food, just like any other customer. But nothing about the moment was ordinary. As the car rolled along with a language model at the controls, one of the engineers cracked a half-joking remark: “Maybe AGI is here after all.” It was a throwaway line meant to capture the absurdity of the moment, but it also carried a deeper point: something significant seems to be shifting in the way AI systems relate to the world around them.
To understand why this moment matters, it helps to remember how self-driving cars have traditionally worked. Autonomous vehicles don’t typically rely on the kind of general-purpose AI model that powers a chatbot. They are built with stacks of specialized software and hardware, carefully trained on millions of miles of driving data, and engineered with redundant systems for every possible edge case. Lidar sensors, radar, high-definition maps, highly specific algorithms for lane detection, pedestrian tracking, and obstacle avoidance. The whole machine is a purpose-built system designed for the task of moving safely through physical space. What the Axiom engineers did was completely different. They didn’t train GPT-6 Astra on driving data. They didn’t spend months tuning it for the challenge of navigating a parking lot. They simply connected a language model to a chat interface, a server, cameras, and the car’s steering system, and let it figure things out in real time. The model was doing something it was never explicitly taught to do. It was looking at the world through cameras, interpreting what it saw, and making decisions based on that information, all without any prior coaching from the trio. That’s a striking leap. It suggests that language models, despite being trained primarily on text, are beginning to pick up a loose but useful understanding of the physical world. They are not just becoming better at words. They are becoming better at space, motion, and even cause and effect. And when a chatbot can use that understanding to pilot a car through a drive-thru lane, it starts to feel like the boundary between the digital and the physical is steadily eroding.
The broader lesson here is that physical reasoning has become one of the great remaining frontiers for artificial intelligence. Today’s most impressive models can answer complex questions, summarize documents, write code, and even act as autonomous agents inside software environments. But step outside the world of screens and servers, and their confidence often collapses. They can struggle with relatively basic human intuitions: knowing that a mug on the edge of a table might fall, understanding that a closed door blocks a path, recognizing that a ball thrown in the air will eventually come back down. These are the kinds of skills that humans learn effortlessly as children, but they remain surprisingly elusive for machines. That may be one reason why, despite the constant claims that AGI has already arrived, many researchers believe physical understanding is still a largely unconquered territory. If artificial general intelligence is supposed to match or exceed human abilities, then it has to be able to navigate the messiness of the real world, not just the tidy abstraction of text. This isn’t just an academic curiosity. A lack of physical reasoning holds AI back in countless practical applications. It’s why robots might be brilliant in controlled factory settings but still awkward and clumsy in a typical home. It’s why digital assistants can order groceries but can’t safely unload them from a car. The gap between having a brain and having a body is enormous, and language models are only beginning to bridge that gap.
Some of the most interesting work on this problem is happening in a wave of new startups, often founded by researchers who left major labs to focus specifically on how AI understands the physical world. Andrew Dai, the CEO of a company called Elorian AI, is a good example. Dai previously worked as a researcher at Google DeepMind, but he left to focus on this exact challenge. His company is betting that better visual reasoning will unlock a whole new class of AI applications. With a stronger grasp of physical scenes, systems could do more than answer questions about images. They could observe a restaurant and understand whether diners are actually enjoying their meals, or watch a living room and anticipate what might happen next. Dai has said that robotics is a crucial test case for these skills. You can’t build robots that meaningfully operate in homes or other unscripted spaces without giving them a solid sense of the physical world. “You can’t really imagine home robotics without this,” he remarked. That sentiment is increasingly shared across the industry. A model that can reason about space and movement is not just a smarter chatbot. It becomes something closer to a foundation for embodied AI, the kind of system that can act in the world rather than just talk about it. To help push the field forward, Elorian and Scale AI, a company known for providing training data to large AI labs, recently created a new benchmark called Humanity’s Sixth Sense. The name is evocative. It suggests something beyond the five senses, a kind of common sense about the physical world that humans possess without even thinking about it. The benchmark is intended to measure how well models can understand physical scenes, testing abilities that go far beyond simple image classification. It’s an attempt to quantify something that feels intuitive to humans but is incredibly hard for machines: the ability to look at a situation and just “get it.”
But as exciting as these developments are, they also raise serious questions about risk and responsibility. There is nothing inherently alarming about an AI model steering a car through a parking lot and up to a fast-food window. It’s a controlled environment, a short distance, and a safety driver was present the entire time. Still, the engineers involved were honest about the stakes. Putting a general-purpose model in charge of a fast-moving, two-ton hunk of steel is a high-stakes undertaking, regardless of how smooth the ride might feel. A language model that has never been specially trained for driving may be able to handle a simple lunch run, but what happens when it encounters an unpredictable situation? A pedestrian stepping into the road unexpectedly, a cyclist weaving between lanes, a child chasing a ball into the street. Human drivers rely on instincts, experience, and a deeply ingrained understanding of risk. A large language model, for all its conversational brilliance, is still a statistical pattern matcher at heart. It can produce decisions that look reasonable, but there’s no guarantee that it truly understands the consequences of those decisions in the same way a human does. As physical understanding improves, AI systems will inevitably be handed more control over real-world actions, not just vehicles, but machinery, medical devices, and perhaps eventually human-like robots. That potential is thrilling, but it is also sobering. The same capabilities that allow an AI to safely navigate a drive-thru could, in a less careful context, become dangerous. The difference might come down to how responsibly these systems are tested, how transparent their limitations are, and how much human oversight remains in the loop.
In the end, the In-N-Out experiment feels like a small but revealing snapshot of where artificial intelligence is headed. It wasn’t a dramatic demonstration of superhuman intelligence. No one was trying to make a political statement about the future of transportation. A few engineers wanted lunch, and they decided to let an AI model drive them there. But in doing so, they showed that language models are beginning to blur the line between words and action. The joke about AGI being here was just a joke, but it touched on something real. Maybe artificial general intelligence won’t arrive as a single, dramatic breakthrough. Maybe it will emerge quietly, through a thousand small moments like this one, when a machine trained on text turns out to have a surprising grasp of the physical world. That’s both inspiring and unsettling. It means the future of AI isn’t confined to screens and servers. It’s coming into the streets, into our homes, into the places where we live and move and eat. The car that rolled up to that take-out window was a test of technology, but it was also a test of trust. Would you let a chatbot drive your car? Would you let a robot make your dinner? The answer may depend on how well these machines continue to understand not just our words, but the world those words describe. For now, the engineers at Axiom can say they made history in the most human way possible: they were hungry, they tried something bold, and they got their burgers. Whether that moment becomes a footnote or a turning point is still up to us. But one thing is clear. The line between artificial intelligence and everyday life is getting thinner, and it may not be too long before we all find ourselves in the passenger seat, wondering who, or what, is really driving.