Enabling technology
No AI Model Recovers What Was Never Captured
Detection is a chain: optics, sensor, edge compute, the AI model, and the operator who acts on the result. Each stage can only work with what the stage before it passed on, so the weakest one sets the ceiling for everything after it. No model recovers detail the optics never resolved; no operator acts on a target the model never saw. That is why the camera is engineered around what the AI has to see, rather than adapted to it afterwards.
Layer 1 · Optics
Aerial Lens Technology
Glass matched to mission altitude and ground resolution — wide-area to standoff — resolving to the full 10K grid, and metrically calibrated so each pixel maps to a known ray.
The optics layer →Layer 2 · Sensor
10K Global-shutter Capture
A 100 MP full-frame sensor at 10,000 × 10,000 pixels, all pixels exposed at once. The global shutter keeps every frame geometrically consistent from a moving platform — the basis for detection and geolocation.
The sensor layer →Layer 3 · Compute
Every Pixel Processed Onboard
Edge compute moves the full 10K stream over PCIe and runs inference on every pixel in real time, within a flight power and thermal budget — so only findings need the downlink.
The compute layer →Layer 4 · Software
Optimise, Infer, Deliver
Your model runs unchanged in an isolated container, fed tiles matched to its input, accelerated for the edge — and findings leave in open standards your systems already read.
The software layer →What engineering all four layers together buys you
- Direct georeferencingOptics + Compute
Metric calibration, IMU and GNSS combine into a camera pose per frame — pixels become target coordinates without a ground workflow.
- Environmental robustnessOptics + Compute
Athermal glass and sealed enclosures hold performance across the altitude and temperature envelope — engineered for the field, not the bench.
- Open architecture · MOSASoftware + Compute
Run the stack on FORGE or on your own NVIDIA hardware — and a trusted, NDAA-compliant supply chain behind all of it.
01 · Optics
The Optics: the Image the Sensor Receives
A sensor can only record the image the lens delivers to it. The optics decide how much of the sensor’s resolution is actually usable, whether that holds as conditions change in flight, and — through careful calibration — whether a pixel can be turned into a real-world direction.
At a Glance
- Precision glass, not plastic, holds sharpness edge-to-edge and across temperature.
- Resolving power is matched to the sensor, so no detail is lost at the lens.
- Infrared-cut filtering keeps colour true for the model.
- Metrically calibrated — every pixel maps to a known angle, the basis for georeferencing.
Glass versus plastic
Lens elements are made from optical glass or from polymer (plastic). Glass is optically superior and far more stable: it expands very little and holds focus as the air cools with altitude and as the camera heats and pressure changes. Plastic is lighter and cheaper, but it shifts with temperature — moving the focus and softening the image just as the environment changes.
Resolving power
A 100 MP sensor only delivers 100 MP of real detail if the lens can draw detail that fine. Every lens has a resolving limit: below a certain size, fine features blur together. If the glass blurs beneath the pixel size, the extra pixels record blur rather than information. Matching the lens’s resolving power to the sensor grid is what makes a high pixel count real rather than nominal.
Infrared-cut filtering
Silicon sensors respond to light well beyond human vision, into the near-infrared. Left unfiltered, that infrared light contaminates colour and softens focus, because it focuses at a slightly different point than visible light. An infrared-cut filter blocks it for accurate colour. Deliberately removing the filter — as on a monochrome path — opens the near-infrared back up for low-light and haze penetration.
Metric calibration
Every real lens distorts the image slightly. Metric — or photogrammetric — calibration measures that distortion and the exact internal geometry, so each pixel corresponds to a precisely known ray leaving the camera. Combined with the platform’s position and orientation, that is what turns a pixel into a coordinate on the ground: georeferencing, computed onboard rather than in a separate ground workflow.
Aerial optics engineered to resolve the full 100 MP grid, athermalised to hold focus across the operating envelope, with controlled IR-cut response and metric calibration as the optical basis for onboard georeferencing. See the ROOK & ECHO sensors →
02 · Sensor
The Sensor: Turning Light Into Data
A digital image sensor is a grid of tiny light-collecting wells — the pixels. How large those wells are, how they are read out, and which light they respond to set a hard ceiling on image quality. Detail that is never captured cleanly cannot be recovered later, so the sensor also sets a ceiling on what an AI model can detect.
At a Glance
- A large-format sensor gathers far more light, and full 12-bit RAW is kept end-to-end — so detail survives in deep shadow and bright highlight where 8-bit pipelines clip.
- An electronic global shutter freezes fast motion with no rolling-shutter skew and no moving parts to wear.
- Passive imaging: the sensor emits nothing, so nothing gives the platform away.
- Resolution and sensitivity are balanced for wide-area detection, not just a pretty picture.
Sensor size and pixel size
A “full-frame” sensor measures 36 × 24 mm — many times the physical area of the small sensors in phones, action cameras and most video cameras. A larger sensor allows larger pixels, and a larger pixel is a bigger bucket: it gathers more photons before it fills up. More photons per pixel means a stronger signal relative to noise and a wider dynamic range — more usable detail in deep shadow and bright highlight at the same time.
Resolution: more pixels, or bigger pixels
Resolution is a trade-off. On a sensor of fixed size, more megapixels means smaller pixels — each collecting less light. Keeping pixels large means fewer of them. A physically larger sensor is what lets you have both: a high pixel count with pixels that are still large. What matters operationally is how many pixels land on the object — its ground sample distance. Too few and a target is an unresolvable blob; enough and it becomes a recognisable shape.
Electronic global shutter
A sensor can be read two ways. A rolling shutter scans line by line, top to bottom, so on a moving platform straight edges skew and fast objects smear. A global shutter exposes every pixel at the same instant, recording one geometrically consistent frame. “Electronic” means it does this with no moving mechanical shutter — nothing to wear or fail — and with very short, precise exposures that freeze motion and reduce blur.
Passive imaging
A camera is a passive sensor: it only receives ambient light. Active sensors — radar (including SAR) or lidar — must emit their own signal to see. That emission is itself detectable and can be jammed. A passive sensor transmits nothing to intercept, though it depends on there being light in the scene to begin with.
Colour versus sensitivity
A colour sensor places a red, green or blue filter over each pixel (a Bayer array), so each pixel sees only part of the spectrum and full colour is reconstructed by interpolation. A monochrome sensor has no colour filters, so every pixel collects the full available light — giving more sensitivity, higher effective resolution, and reach into the near-infrared, which helps in low light and haze. The trade is colour information against raw light-gathering.
A 100 MP full-frame sensor (10,000 × 10,000 px) in RGB and monochrome variants — the monochrome path extending into near-infrared — with an electronic global shutter, operating fully passively. See the ROOK & ECHO sensors →
03 · Compute
The Compute: Why the AI Runs on the Aircraft
Where the AI runs is itself an enabling choice. Running it on the aircraft — at the edge — rather than sending imagery to a ground station or the cloud changes what is possible, and it is only possible because the compute is engineered to fit a flying platform.
At a Glance
- The full 10K stream reaches the processor over a wide internal PCIe bus — nothing dropped or compressed.
- Every pixel is processed onboard; only findings leave the aircraft, never raw video.
- Decisions happen in the moment — no ground round-trip.
- Serious AI in a sensor’s power budget — and it keeps working through a link blackout.
Moving the data: a high-bandwidth internal bus
A 10K sensor produces an enormous stream of data every second — far more than an ordinary camera connection can carry. PCI Express (PCIe) is a very wide, high-speed data path used inside computers; here it links the sensor directly to the processing module so the entire full-resolution stream reaches the GPU without a bottleneck. It is a separate internal path — not the radio link to the ground — and it is what makes onboard, full-frame processing possible in the first place.
Process onboard, not on the ground
A 10K video stream is far too much data to send over a radio link. Transmitting raw imagery is slow, floods the link, and stops entirely if the link drops. Processing on the aircraft means only the findings — a handful of small messages — ever need to be sent.
Answers in the moment
Sending imagery to the ground, waiting for it to be processed, and getting a result back all takes time. Deciding onboard removes that round trip — the answer is available while it still matters, not seconds or minutes later.
A data-centre’s job in a sensor’s power budget
An aircraft offers very little power and no room for fans. Edge compute has to deliver serious AI performance within a few watts, in a small sealed package that sheds heat passively — and keep doing it as the camera warms up. Size, weight and power (together, SWaP) set the ceiling on how much AI can actually fly.
Keep everything; keep working
Fast onboard storage retains every raw frame for detailed review after the flight, while only findings are sent live. And because the processing is self-contained, the mission keeps working even if the communications link is jammed or lost — findings simply queue until it returns.
Many sensors, one clock
When more than one sensor flies on the same platform — several heads for a wider swath, or a mix of sensors — their frames only combine cleanly if they are captured at the same instant. Sapient synchronises capture across sensors to a shared timebase, so every frame shares a common trigger and timestamp rather than drifting apart from one exposure to the next.
NVIDIA Jetson Orin NX (16 GB) edge compute on a Sapient-designed carrier board — with a 2 TB NVMe SSD for full-mission RAW retention, flight interfaces, and a sealed 160 g package (190 × 55 × 30 mm) that mounts anywhere. See the FORGE module →
04 · Software
The Software: From Frame to Finding
Hardware captures the image; software turns it into an answer. The enabling ideas here are less visible than glass and silicon, but they decide whether an AI model can run on the aircraft at all, how well it performs, and whether its output is usable the moment it lands.
At a Glance
- Your model runs unchanged inside an isolated, encrypted container.
- Inputs are shaped to exactly what the model expects — scaled, tiled, calibrated.
- Fast enough to keep pace with the sensor in real time.
- Outputs speak the receiver’s language — Cursor-on-Target (CoT), straight into ATAK and C2.
Containers: a sealed box for the model
A container packages an application together with everything it needs to run, so it behaves identically wherever it runs and cannot reach anything outside its box — like a shipping container: standard on the outside, your cargo sealed inside. Each AI model runs in its own container. You bring a model as-is, updating it means swapping the box rather than rebuilding the system, and because the box is sealed the model and its data stay isolated.
Feeding the model what it expects
An AI model is trained on images of a fixed size and a particular scale — how much ground each pixel covers. A 10K frame is far larger than that, and its scale changes with altitude. The software cuts each frame into tiles the size the model was trained on, and resamples them to the scale it learned, so every tile looks familiar. This is possible because the optics are calibrated and each frame carries its own pose, so the ground scale of every pixel is known.
Fast enough for real time
Running a large model on every frame, in the air, on a few watts, means tuning the model to the exact onboard processor — using lower-precision arithmetic and fused operations that do the same work in far fewer steps. This is the difference between analysing a live stream and steadily falling behind it.
Speaking the receiver’s language
A finding is only useful if the systems downstream can read it. Rather than inventing a new format, the software emits findings in the open message standards and file types command systems already use — so nothing new has to be installed on the ground to receive them.
Models run in isolated Docker containers, accelerated for the edge with TensorRT, with native-resolution crops kept in the open RAW DNG format, and findings delivered as Cursor-on-Target (CoT) into ATAK and existing C2. See the IGNITE:AI framework →
Designed together
Capabilities No Single Layer Could Deliver
The defining capabilities of the platform sit between the layers — they exist because optics, sensor, compute and software are engineered as one system.
See Why the Engineering Matters
Explore sets out what these technologies buy you in the field — coverage, altitude, and a near-silent RF footprint.