Re-imagining the scope of Physical AI

Physical AI as defined currently as “artificial intelligence systems that perceive, reason about, and act directly within the physical world, rather than existing purely as software or behind a digital screen. While traditional AI (like large language models or image generators) works with information to output text, code, or pixels, Physical AI connects digital intelligence to physical hardware—bringing the power of machine learning from the realm of “bits” into the realm of “atoms”. Physical AI systems are “adaptable”, learn from their environments in a continuous feedback loop. They perform three core tasks – sensing (using “sensors” to learn about the environment), reasoning – understand – spatial relationships, geometry and laws of physics, and finally action in the environment – modulating the physical world – including motion, manipulation and more. Given the above definition, much of Physical AI R&D is focused on building robots, and intelligent hardware, autonomous vehicles for different environments. We believe this is too narrow an interpretation of Physical AI and we are missing addressing a larger variety of problems where physical world knowledge is essential for decision-making but need not be embedded in a machine or a robot. Shown below is a grid defined by: a) The nature/size of the physical space on the horizontal axis and b) the “number” of agents that occupy these physical spaces and are actively involved in sensing, reasoning and action in these physical spaces (both human and autonomous). Each grid cell outlines potential domain/problem-solving tasks that these agents perform in that space. The grid illustrates the wide variety of tasks physical AI needs to address involving spatial reasoning, geometry, physical world understanding and more.

However, current physical AI efforts focus on the top LHS cell (single agent/local space). Autonomous vehicle efforts focus on a single driver agent – a collection of autonomous vehicles work – because of an “implicit” environment or system agent – that monitors/enforces the traffic rules and each of the individual agents also voluntarily subscribe to these rules. More on this implicit “system agent” below. If human drivers and autonomous vehicles do not follow traffic rules, the whole system would fail (as in SE Asian countries!). Much of the spatial foundation model building (such as those from Google Earth and more) also has focused on the bottom RHS – combing remote sensing data for various tasks – where the number of agents is too large and one is only studying collective effects on the region of interest. The highlighted cells in the grid below are “complex physical AI related” problems which are not yet fully under the rubric of Physical AI though many aspects overlap. We characterize these problems and outline the similarities between these problems on how they need a common spatial representational framework in this essay.

For completeness, a comparison between foundational models built for the top LHS and bottom RHS cells is shown below –

It is unclear what kind of “Foundational Models” will address the problems in the rest of the cells in above table. Is even a Foundational Model a part of the appropriate solution and is it going to power solutions that are useful in answering the relevant questions in those problem domains? Before further discussion on the representational issues, we expand on the core idea behind the grid – the different dimensions of each cell. We only picked two essential ones but there are a few additional ones that characterize these problems. The dimensions of each cell include :

a) The size of the physical area under consideration and its “structure”, “properties” etc. (used above)
b) The number of agents – both human and autonomous in these spaces. (used above)
c) The capabilities of these agents – Are these agents heterogeneous or homogeneous (different or same capabilities)? Are all agents in a given physical space independent decision-makers or is there a “command structure”?
d) Is there a “system” or environment agent – that is central and coordinates the groups of agents? It can be explicit or implicit in the environment or the protocols of coordination? A lot of our real-world with agents works because “intelligence” is embedded in the environment explicitly or implicitly in rules of coordination and behaviors. If agents violate these rules, the system descends into chaos. For example, a Waymo vehicle relies on the intelligence embedded in the environment – traffic signs, road markers – yellow lines, while lines and traffic signals etc. and assumes others on the road would follow it too.
e) The different type of entities that occupy a given space – what kind of world knowledge does each agent need to know – hydrants, parked vehicles, trees, shrubs, people, garbage on the road etc in the context of driving.
f) The types of tasks each individual agent is capable of doing including different types of agent skills and task complexity. This includes agents that need to deal with both the digital and physical world.
g) The rate at which these “physical spaces” evolve dynamically and how quickly can the agents sense/perceive their environments and update their incoming data rates and external action frequencies. What can be assumed to be changing slowly or quickly in the realworld? How does an agent detect the change unless one is looking for the change? For example, in the context of driving, what if the road gets flooded while driving.
h) Time horizon and type of agent decisions – short term versus long term, goal oriented versus reactive, periodic versus opportunistic
i) How much does one agent need to know (or assume) about the “mental models” and “capabilities” of the other agents? Are the agents working co-operatively or adversarily? How does the agent update its own mental model of other agents in its physical environment?

Among all the dimensions above – an understanding of the “physical world” around an agent or group of agents is essential. This dimension critically interacts with all the other dimensions in different ways – simplifying a problem for an agent or making it extremely complex. Imagine a Roomba cleaning a home versus a Roomba-like autonomous road sweeper – they need very different models of the physical world and agents embedded in them to function effectively though they may share the same perceive, reason, and react/act loop. The R&D on geo-spatio-temporal reasoning for solving problems in the above grid spans multiple themes – a) GeoAI, b) GeoSpatial Foundational models, c) Spatial and SpatioTemporal reasoning, d) Physical world modeling, e) Spatial science, f) Location intelligence g) PINNs, and h) Vision LLMs. We believe all the above problems share a number of common issues that need to be addressed just from a spatial, geometric and topological representation and reasoning perspective which we outline below. Discussion on how these models may be used to compute via agentic technology is discussed in the essay on agent technologies.

All the different spatial reasoning related problems at different resolutions share the following representational issues that need to be addressed:

  1. Geo-spatial phenomena data is dominated by – Spatial auto-correlation, modality imbalances (image data dominates in contrast to other sensor types), inconsistent and non-standard cross-modality fusion, temporal discrepancies – different data sets at different sampling rates, spatial heterogeneity and non-stationarity.
  2. Geo-spatial phenomena occur concurrently at multiple-scales – how to capture and combine, how to aggregate/disaggregate properties across different resolutions
  3. Geo-spatial phenomena exhibit cause-effect relationships that have lags/delayed effects. Long horizon events have complex long running causal chains.
  4. Geo-data representations are inconsistent -Pixel (aka point data and point clouds) versus understanding of topological relationships. Currently no deep spatial understanding exists in FMs, FMs lack understanding of physical laws, retrieve “patterns” memorized or hallucinate, lack understanding of 2D, 3D and 4D (with time) spatial phenomena understanding.
  5. Spatial digital representation across two mutually inconsistent formats – vector and raster. Reconciling these two formats into one standard “model” is a difficult issue. Further coordinate systems across data sets may not line up. FMs lack a definitive understanding of absolute and relative coordinates while reasoning with geo-spatial relationships.
  6. Geographical data sets are biased with geographies where data is available – not representative of different parts of the world.
  7. Furthemore, we need to reconcile and standardize discrete and continuous geo-spatial models and data in a consistent framework. We need to the ability to sample, interpolate and extrapolate between these two basic data representations.
  8. Finally, all LLM stacks are not grounded in “physical reality”, all vision models cannot resolve evolving topologies and geometries that are driven by physical laws.

Addressing the above representational issues is essential to build a reliable model (or a collection of loosely coupled models) that can offer consistent explanations for various geo-spatial phenomena (data co-relationally or causally). Furthermore, building one or more embeddings, and combining them in consistent ways is a difficult problems. Embeddings can only answer “similarity” based queries possibly of different kinds, however access to component data sets is essential to drive problem-centric reasoning. Validating a geo-spatial embedding at scale is a difficult task.

Resolving the aforementioned representational models is closely coupled with the “mathematical models” that are used to compute responses to different types of queries. An API for such a system – should at least do the following –

  • Provide access to different types of “component” data and its set of attributes
  • Provide a “component’s” representation in an embedding – for example how is a city center represented versus a suburb (what are the vector coordinates)
  • Provide facility to compute similarities (along different attributes – ) and visualize relationships
  • Provide the ability to run “compute models” to infer properties – such as aggregating a property across regions, disaggregating a property given a representation scheme, converting a property from its H3 representation to a S3 or geohash representation, approximating properties for validation/verification.

An API which provides such support will faciliate geospatial data scientists to validate and verify the model – computationally since there is no one size fits all embedding. Providing these basic constructs then helps one layer in the domain workflows for tasks such as Site-evaluation, trade area analysis and more given additional business data. Once a reasonable representation addresses some of the above issues – the following engineering issues that need to be sorted at scale to build out and maintain the system are:

  1. Scaling FM building including all embeddings as the geo data evolves and changes with time. How do you find and update the deltas?
  2. Resolving all the data consistency issues including biases and privacy. How do you check this?
  3. Building benchmarks and training data sets for answering different types of geo-spatial queries that stress the representation and validating the responses with ground reality. Lot of the current benchmarks are highly ad hoc.
  4. Addressing the big-data combinatorics across multiple ways of organizing, indexing spatial data at different granularities.

Utilizing the above model(s) in different “agent” architecture configurations (involving humans, software agents and robots) is more an engineering issue rather than a fundamental modeling one. Addressing the afore-mentioned geo-representational issues in a consistent framework is a major first step towards realizing the impact of AI technologies in Geo-spatial reasoning. Considering every one is building FMs at different scales of resolution, what would it take to build a Geo-Spatial Verification Service? Something to seriously think about.