Machine vision has spent years moving in one direction: more.
More pixels. Higher frame rates. More powerful processors. More data flowing from the sensor into memory and through increasingly sophisticated algorithms.
But much of what a camera records from one frame to the next hasn’t changed.
A 4K camera looking down a warehouse aisle captures more than 8 million pixels. A fraction of a second later, it captures them all again. The walls are still there. The floor is still there. Most of the equipment hasn’t moved. Yet all of those pixels become data that has to be read from the sensor, transferred, stored and potentially processed before a vision system decides what matters.
Researchers at Brown University recently demonstrated what happens when a machine-vision system starts with a different assumption.
Instead of recording everything and sorting it out later, their experimental system records primarily what changes.
The approach combines an event-based camera with a spiking neural network, creating a vision pipeline in which both the sensor and the processor work with sparse events rather than a continuous stream of complete images. In laboratory experiments, the system tracked and reconstructed moving objects through dense fog and turbid water that obscured the targets from a conventional camera.
The fog makes for a striking demonstration. But the more interesting engineering story is what the researchers chose not to record.
A Camera That Doesn’t Take Pictures
A conventional digital camera captures frames.
At a set interval, the image sensor measures the light reaching its pixels and creates another complete image. At 30 frames per second, that happens 30 times every second whether the scene changed dramatically or barely changed at all.
An event camera, also called a dynamic vision sensor (DVS), operates differently.
Its pixels work independently. Rather than waiting for the next frame, each pixel monitors the intensity of the light reaching it. When that intensity changes beyond a set threshold, the pixel generates an event containing information about what changed and when.
If nothing changes, the pixel stays quiet.
There isn’t a conventional sequence of pictures coming from the sensor. Instead, the output is a stream of asynchronous events representing changes across the scene.
That makes event cameras particularly good at detecting motion. The sensor doesn’t need to wait for one frame to end and another to begin before recognizing that something moved. Changes are reported as they happen.
It also means large portions of an unchanging scene generate little or no data.
For a machine trying to determine where an object is going, that might be exactly what it needs.
Brown researchers Ning Zhang and Arto Nurmikko took advantage of this behavior to tackle an imaging problem that becomes particularly difficult for conventional cameras: seeing through scattering environments.
Let the Sensor Ignore the Fog
Fog doesn’t simply make an image darker. Water droplets scatter light traveling between an object and a camera.
Some of the light reflected from an object still reaches the sensor, but it has been mixed with scattered light. As scattering increases, the spatial information needed to distinguish the object becomes increasingly difficult to recover.
Murky water creates a similar problem as suspended particles scatter incoming light.
There are several engineering approaches to seeing in those conditions. LiDAR, radar, infrared imaging, time-of-flight techniques and computational imaging each provide different ways of extracting information when conventional visible-light imaging begins to fail.
The Brown team approached the problem using motion.
A moving object creates relatively rapid changes in the light reaching individual pixels. The scattering background changes more slowly.
An event sensor is already designed to look for exactly that difference.
Instead of repeatedly capturing the target along with the fog surrounding it, the dynamic vision sensor responds primarily to the changes associated with the moving object. Much of the slowly varying scattering background generates fewer events.
The sensor effectively begins filtering the scene before the information ever reaches the processor.
In experiments, the researchers sent images of alphanumeric characters and bird silhouettes through a fog chamber and through turbid water. The targets followed random trajectories rather than moving along predetermined paths.
A standard camera eventually lost the objects in the scattering environment. The event sensor continued producing information about their movement.
There was one problem.
Those events didn’t look much like an image.
When the Sensor and Processor Speak the Same Language
The raw output of an event camera is closer to a cloud of activity than a photograph.
Individual pixels report increases or decreases in brightness along with precise timing information. Somehow, those scattered events have to become useful information about what the object is and where it is going.
Brown’s researchers paired the sensor with a deep spiking neural network (SNN).
That pairing matters.
Most artificial neural networks operate on numerical values passed through layers of artificial neurons. Image-processing networks typically receive arrays of pixel values representing complete images or batches of images.
A spiking neural network works differently. Information is represented by discrete spikes occurring over time.
The concept is loosely inspired by biological neurons, which communicate through electrical impulses. A neuron remains relatively quiet until its internal state reaches a threshold, at which point it fires.
That means the Brown system has an unusual relationship between its sensor and its processor.
The camera produces asynchronous spikes.
The neural network processes asynchronous spikes.
There is no need to turn the event stream into a conventional video feed before processing it.
The researchers designed the SNN with two parallel jobs. One module estimates the location of the moving object while another reconstructs its shape. Leaky integrate-and-fire neurons within the network also retain temporal information, allowing the system to accumulate evidence from events occurring across multiple time steps.
In the experiments, the reconstructed images achieved structural similarity scores ranging from 0.81 to 0.96, depending on the test conditions. Full-sequence inference was completed in less than 15 milliseconds.
The result wasn’t a detailed photograph. It was a reconstruction of the target’s shape paired with information about where it was moving.
For a machine, that distinction could matter more than image quality.
The Cost of Seeing Everything
A person watching a video generally wants the complete picture.
A robot might not.
Consider an autonomous mobile robot moving through a warehouse. Its control system might need to know that a worker stepped into its path, where the worker is and how quickly the distance between them is closing.
It doesn’t necessarily need another high-resolution photograph of the wall behind that worker.
The same problem appears in drones, autonomous vehicles, industrial machine vision, and other edge systems.
Increasing sensor resolution generates more data. Increasing frame rate generates it more frequently. That data then has to move from the image sensor to memory and eventually into a processor, where algorithms determine which parts are important.
Processing isn’t the only energy cost. Moving and storing data also consumes power.
That becomes increasingly important in systems where energy, bandwidth and latency are limited.
An event camera changes the problem at the beginning of the signal chain. If a pixel has nothing new to report, it produces no event.
That means the system doesn’t have to capture an enormous amount of redundant visual information and discard it later. Some of that filtering happens inside the sensor itself.
The Brown team’s spiking neural network continues the same philosophy on the processing side. Computation is driven by sparse events rather than dense image frames.
In their analysis, the researchers estimated that the SNN required about 18 times less computational energy than an equivalent conventional artificial neural network.
The dynamic vision sensor itself operates at low power, with Brown reporting consumption in the tens of milliwatts.
That doesn’t mean an event camera paired with an SNN automatically makes every machine-vision system more efficient. Hardware implementation, memory architecture, workload and the density of events all matter.
But it points toward a different way of designing the sensing pipeline.
Instead of collecting everything and relying on a powerful processor to decide what is useful, the sensor itself becomes part of the decision about what information deserves to move farther into the system.
What Gets Lost When You Record Less?
Recording only changes comes with an obvious tradeoff.
Something has to change.
A stationary object produces little useful event information because the brightness reaching the relevant pixels isn’t changing enough to trigger them.
That creates a serious limitation for many of the applications the technology might eventually serve.
An autonomous vehicle needs to detect a stopped car just as much as a moving one. A search-and-rescue drone cannot ignore someone who isn’t moving. An underwater robot still needs to know about a stationary rock.
The current Brown system also reconstructs silhouettes rather than complete grayscale images, and its sensitivity declines under very low-light conditions.
Those limitations make it unlikely that an event camera would simply replace every conventional camera in a machine-vision system.
A more realistic direction could involve combining different sensing technologies.
A conventional camera could provide detailed visual information. An event camera could handle rapid motion and changes. Radar, LiDAR or time-of-flight sensors could contribute distance and depth information.
Brown’s researchers are already investigating ways to expand what their system sees. One possibility is a light-intensifier front end for darker environments. They are also exploring depth-resolved approaches, including time-of-flight measurements and stereo event cameras that could recover information about three-dimensional targets.
The important question isn’t necessarily whether an event camera is better than a conventional camera.
It’s what information each sensor should be responsible for collecting.
Rethinking What a Machine Needs to See
The Brown system is interesting because it succeeds without producing the kind of video we normally associate with a camera.
It doesn’t see through fog by creating a clearer conventional photograph and then handing it to an image-processing algorithm.
It changes the sensing problem first.
Moving objects create events. Much of the slowly changing scattering background does not. Those events move directly into a processor designed to work with the same sparse representation.
The result is a vision pipeline built around change rather than images.
That idea extends well beyond fog.
As engineers put increasingly capable vision systems into robots, vehicles, drones and industrial equipment, the amount of sensor data those systems generate continues to grow. Faster processors help deal with that information, but they don’t address the question sitting farther upstream.
Does the machine need all of it?
A person looking through a camera wants a picture. A robot trying to avoid an obstacle might need something simpler: Something moved. It’s here. It’s heading there.
If that’s all the machine needs to know, capturing everything else first might be the inefficient part of the design.
The Research
Brown University, October 1, 2026
Brown University engineers build brain-inspired imaging system that can ‘see’ through fog
Primary university release covering the system, experiments, potential applications and comments from Ning Zhang and Arto Nurmikko.
Brown University research announcement
Ning Zhang and Arto Nurmikko, Advanced Science, August 24, 2026
Neuromorphic Optical Tracking and Imaging of Randomly Moving Targets Through Dynamic Dense Scattering Media
Original peer-reviewed paper detailing the DVS, spiking neural network architecture, experiments, latency, reconstruction results and energy analysis.
Advanced Science research paper