Or: when pixels meet atoms

In recent months, the discussion around artificial intelligence has revolved almost exclusively around large language models (LLMs) and generative chatbots. But while the world experiments with text prompts, we at IoT are working on a different frontier of AI: the interface between digital intelligence and the unstructured complexity of our real world.

Our current research project shows that the concrete value of AI is often to be found outside the digital world – namely where pixels meet atoms.

The challenge: complexity in physical space

In many sectors, manual control and data-capture processes are still the standard. These analog procedures are not only extremely resource-intensive but often have a surprisingly high error rate. 20% errors and more are not uncommon.

Whether it’s the precise recording of inventory, quality control in production, or the verification of complex physical layouts – people reach the limits of their concentration with repetitive, high-frequency tasks.

The technological approach: spatial intelligence & edge-native AI

To close this “analog gap,” we are developing an integrated overall solution that goes far beyond simple image recognition. Our approach combines three technological pillars:

  1. Edge computing & real-time processing: All computing power is provided directly on the end device (smartphone or tablet). Through methods such as quantization and pruning, we optimize complex models for local use. This guarantees real-time processing that works completely independently of cloud connections or unstable networks.
  2. Computer vision (CV) at a state-of-the-art level: We rely on a combination of highly effective architectures. While convolutional neural networks (CNNs) such as ResNet or EfficientNet form the basis, for high-precision tasks in complex scenarios we use YOLO, Faster R-CNN, or vision transformers (ViT). These models not only identify objects but link them directly with attributes such as brands or product properties.
  3. Spatial computing via AR & SLAM: A purely 2D analysis is often not enough in unstructured environments. By integrating augmented reality (AR) and SLAM algorithms (Simultaneous Localization and Mapping), we capture the exact position of objects in three-dimensional space. This creates a virtual matrix that recognizes overlaps and rules out double counting – for example, with movable displays.

Impact: precision as the new standard

With this automated end-to-end solution, we transform error-prone processes into high-precision workflows:

  • Error minimization: The rate drops from 20% to a maximum of 2%.
  • Efficiency increase: Processing time is reduced by up to 50%.
  • Sustainability through data accuracy: More precise inventory data reduces goods losses by an estimated 30% and avoids unnecessary rework, which has an immediate effect within the company in times of high labor costs and a shortage of workers.

A universal blueprint

Although our current validation takes place in a specific application scenario, the technological core is universally scalable. The combination of robust edge AI and spatial AR recognition serves as a model for future solutions in logistics, industrial quality control, traffic management, or healthcare.

Technology with its feet on the ground

We at IoT have always placed people at the center of digital transformation. As fascinating as the technical foundations and solutions may be, if they don’t work in the real world, they are worthless.

We don’t build chatbots. We build assistants that understand the real world, in order to relieve people of repetitive tasks and create space for value-adding activities.

Stay tuned.