A generation ago, an embedded product often had one brain.
It was not always a powerful brain. It did not need to be. A microcontroller read a few inputs, made a few decisions, drove a display or motor and repeated the cycle. The software fit in modest memory. The timing was predictable because the machine had little else to do.
Those systems still exist. A smoke detector does not need a Linux distribution. A simple thermostat does not need a neural-processing unit. But once a product must recognize people, show a rich interface, connect securely to a network, update itself in the field and respond to a physical event on time, the old model starts to fray.
The processor running the user interface is not necessarily the processor that should stop a motor. The processor handling a camera stream is not necessarily the one that should manage secure boot. The core best suited to machine-learning inference is rarely the right place for an Ethernet stack, a touch display and a high-priority control loop.
So embedded systems are dividing the work.
The division does not always involve three separate chips. More often, it occurs inside one system-on-chip or between an applications processor, a microcontroller and a purpose-built accelerator. The arrangement is becoming familiar: one processor runs the operating system and the human-facing parts of the product; another deals with deterministic, real-time work; a third handles the computationally repetitive task, such as vision inference, signal processing or a high-speed hardware interface.
It is a useful way to understand a modern robot, medical device, industrial camera, EV charger or connected appliance. One product, three processors, three different ideas of what matters.
The old division of labor never really disappeared
Specialized processing is not new. It predates most of the products now described as edge devices.
Texas Instruments introduced the TMS1000 microcontroller family in 1974, when a single chip able to handle control work changed the economics of everything from calculators to appliances. [1] Around the same time, another type of processor was taking shape. Digital signal processors were built around a simple insight: handling streams of audio, radar, communications or sensor data required a different kind of arithmetic machine from the one running general-purpose program logic.
The first single-chip DSPs appeared in the late 1970s. TI’s programmable TMS320 family arrived in 1983 and found its way into products ranging from toys to communications equipment. [2] By the early 2000s, the boundaries had begun to blur. TI was shipping system-level DSPs that combined a TMS320C5000 DSP with an ARM7 processor, arguing that designers no longer needed to buy separate chips for control, signal processing and an embedded operating system. [3]
The hardware has changed beyond recognition since then. The architectural argument has not.
A DSP excels at repetitive numerical work. A microcontroller is built for deterministic response to physical inputs and outputs. An applications processor has the memory, operating-system support and raw flexibility for networking, graphics, databases and product software. Engineers once placed those roles on different boards or different chips. Today, they increasingly sit together in a package, sometimes alongside an FPGA or neural-processing engine.
Integration solved many board-space and power problems. It did not erase the need to decide which work belongs where.
The processor for the screen should not be responsible for the brake
Take an industrial mobile robot moving through a warehouse.
It needs to run a user interface, communicate over Ethernet or Wi-Fi, maintain maps, coordinate with a fleet-management system and log data for service technicians. Those are comfortable jobs for an applications processor running Linux. Linux brings an ecosystem of drivers, networking tools, security updates and application frameworks. It also brings variability. A scheduler may delay a task. A storage write may arrive at the wrong moment. A background service may consume resources nobody expected during the first lab demonstration.
None of those characteristics make Linux bad. They make it the wrong place for the most time-sensitive loop in the machine.
The robot’s motor-control loop needs a reliable response at a known interval. A safety sensor needs attention when it changes state. A power stage does not care whether the product’s web dashboard is refreshing. Those jobs fit a real-time microcontroller core or a dedicated control processor, often running an RTOS or bare-metal firmware.
Then there is vision. A camera mounted on the robot may produce more information than either general-purpose core wants to sift through frame by frame. Image preprocessing, neural-network inference or depth estimation belong on a DSP, GPU, NPU or FPGA fabric, depending on the latency, power and determinism required.
Now the product has three processor domains:
- An applications processor for Linux, networking, data logging and the interface
- A real-time controller for time-critical control and I/O
- An accelerator for vision, inference or signal-heavy workloads
The arrangement is not an indulgence. It is often the only sensible way to keep a sophisticated product responsive without letting one software problem spill into every other function.
STMicroelectronics’ STM32MP2 family offers a compact illustration of the idea. Its parts pair Arm Cortex-A35 application cores with a Cortex-M33 microcontroller core and, on selected models, an NPU. ST describes the M33 as a trusted domain able to isolate critical resources and manage startup and reset of the A35 processors. Its smaller Cortex-M0+ domain is intended to keep selected peripherals active at low power while other cores are stopped. [4]
That is more than a list of cores. It is an operating model. The product is being separated into functions with different timing, security and power requirements.
A fast chip is not a complete architecture
Marketing often turns heterogeneous processing into a cores-and-TOPS contest. Six application cores look better than four. A neural accelerator with a larger number looks better than one with a smaller number. Those figures matter, but they rarely explain why the chip belongs in a product.
The more useful question is: what work must continue when the rest of the system is busy, asleep, rebooting or under attack?

NXP’s i.MX 95 makes the answer visible. The device combines six Arm Cortex-A55 cores, one Cortex-M7, one Cortex-M33 and an eIQ Neutron NPU. It also includes connectivity designed for products that need networking, timing synchronization and industrial interfaces. [5] An engineer designing an industrial vision system could use the A55 cluster for the operating system and applications, the M7 for bounded real-time tasks, the M33 for security-oriented management and the NPU for model inference.
But having the pieces does not automatically produce a good system.
A camera image must get from its input to the inference engine. The inference result must reach the control software. The control decision must travel to the real-time loop. Each handoff introduces questions about memory ownership, latency, error handling and version compatibility. A missed message inside a multi-processor product can be harder to diagnose than a missed wire on a single-board controller.
That is why the software architecture matters as much as the silicon.
TI’s AM62A software-development kit, for example, includes Linux support, RTOS support on its Cortex-R5F core and examples for interprocessor communication between heterogeneous on-chip cores. [6] The presence of those examples says something important. The difficult work begins after the block diagram. The team has to establish which processor owns each peripheral, what happens when one processor resets and how the others continue safely.
Engineers who have worked through those decisions tend to become suspicious of diagrams where every arrow points neatly toward “AI.” Real systems have more mundane problems: shared-memory corruption, incompatible caches, a boot sequence in the wrong order or two teams assuming they own the same GPIO.
The third processor does not have to be an NPU
Neural processing units get the attention because they fit neatly into the story of edge AI. They are not the only third processor in town.
In a high-speed inspection system, FPGA logic might receive and preprocess camera data while an applications processor handles the network and a microcontroller manages deterministic timing. In a software-defined radio, a DSP might perform filtering and modulation while the CPU manages protocol stacks and configuration. In a motor drive, dedicated PWM and control hardware might take over work that a general-purpose core would handle too slowly or too unpredictably.
Microchip’s PolarFire SoC family shows another version of the split. The devices combine a five-core RISC-V cluster with FPGA fabric. The processing subsystem includes one monitor core and four application cores, while the fabric provides a place for custom logic, deterministic I/O and workload-specific acceleration. Microchip positions the family for Linux and real-time applications in the same device. [7]
This matters in products where the input arrives too quickly, the timing is too tight or the interface is too unusual for a software-only pipeline. Engineers do not reach for programmable logic because they enjoy adding another toolchain. They reach for it because, at some point, software running on a CPU is no longer the cleanest way to make the machine behave.
A good architecture does not distribute work because it has available blocks. It distributes work because each block earns its keep.
The hard part is drawing the borders
A three-processor product introduces a management problem. Someone needs to define the borders.
What belongs on Linux? What must remain available when Linux is rebooting? What does the real-time core do if the AI model crashes? Which processor receives a safety-related sensor input first? Who is allowed to update firmware on the accelerator? Which memory is shared and which memory is protected?
Those questions reach beyond software. They affect board design, power sequencing, reset circuitry, debug access, functional safety and field service.
The Cortex-R5 core, widely used in deeply embedded real-time systems, was designed with error management and functional-safety support in mind. [8] That makes it attractive for work where predictable response matters. Yet moving a task onto an R-class core does not make the whole product safe. The interfaces surrounding it still need to be designed, tested and documented.
The same goes for low-power operation. A product that leaves its applications processor awake to watch for a button press or sensor threshold is wasting more than energy. It is giving the most complex processor in the system responsibility for a job a smaller domain could handle while the rest of the product sleeps.
The best partitioning decisions usually come from looking at failure first.
If the network stack locks up, what must still work?
If an over-the-air update fails, which processor still has authority to recover the product?
If the vision algorithm takes longer than expected, what prevents a robot or machine from acting on stale information?
If the NPU is powered down, what basic function remains?
Those are not questions reserved for autonomous vehicles or surgical robots. They belong in the first architectural meeting for any connected device with a physical job to do.
The return of systems engineering
For years, embedded development rewarded integration. Put more capability on one chip, shrink the board, reduce the bill of materials and move on. The next phase rewards partitioning.
Not separation for its own sake. Separation with a reason.
The applications processor should be free to run a modern interface, maintain secure connections and take advantage of a rich software ecosystem. The real-time controller should be free to watch the physical world without waiting for a graphical update or cloud transaction. The accelerator should be free to process the workload it was built for without carrying an entire operating system around it.
That is not a retreat from integration. It is what mature integration looks like.
The product still arrives as one device. The customer still sees one machine. Underneath, however, it behaves less like a single computer and more like a small team with sharply defined jobs. One member talks to the outside world. One keeps time. One does the heavy lifting.
For embedded engineers, the challenge is no longer finding a processor powerful enough to run the product. It is deciding which processor deserves each responsibility, then making sure the handoffs are as dependable as the code running on either side.
Sources
- Computer History Museum, “1974: General-Purpose Microcontroller Family Is Announced.” Computer History Museum
- Computer History Museum, “1979: Single-Chip Digital Signal Processor Introduced.” Computer History Museum
- Texas Instruments, “TI Delivers Two System-Level DSPs Integrating a DSP, RISC Processor and Embedded Operating-System Support.” ti.com
- STMicroelectronics, “STM32MP2 Microprocessor Series.” STMicroelectronics
- NXP Semiconductors, “i.MX 95 Applications Processor Family.” nxp.com
- Texas Instruments, “PROCESSOR-SDK-AM62A.” TI.com
- Microchip Technology, “PolarFire SoC FPGAs.” Microchip Technology
- Arm, “Cortex-R5.” Advanced Error Management & SoC Integration