Enabling technology

No AI Model Recovers What Was Never Captured

Detection is a chain: optics, sensor, edge compute, the AI model, and the operator who acts on the result. Each stage can only work with what the stage before it passed on, so the weakest one sets the ceiling for everything after it. No model recovers detail the optics never resolved; no operator acts on a target the model never saw. That is why the camera is engineered around what the AI has to see, rather than adapted to it afterwards.

RESOLVED · TRACKEDSAME SCENE · SAME MODELDETAIL THE OPTICS NEVER PASSED ON

Layer 1 · Optics

Aerial Lens Technology

Glass matched to mission altitude and ground resolution — wide-area to standoff — resolving to the full 10K grid, and metrically calibrated so each pixel maps to a known ray.

The optics layer →

Layer 2 · Sensor

10K Global-shutter Capture

A 100 MP full-frame sensor at 10,000 × 10,000 pixels, all pixels exposed at once. The global shutter keeps every frame geometrically consistent from a moving platform — the basis for detection and geolocation.

The sensor layer →

Layer 3 · Compute

Every Pixel Processed Onboard

Edge compute moves the full 10K stream over PCIe and runs inference on every pixel in real time, within a flight power and thermal budget — so only findings need the downlink.

The compute layer →

Layer 4 · Software

Optimise, Infer, Deliver

Your model runs unchanged in an isolated container, fed tiles matched to its input, accelerated for the edge — and findings leave in open standards your systems already read.

The software layer →

What engineering all four layers together buys you

  • Direct georeferencingOptics + Compute

    Metric calibration, IMU and GNSS combine into a camera pose per frame — pixels become target coordinates without a ground workflow.

  • Environmental robustnessOptics + Compute

    Athermal glass and sealed enclosures hold performance across the altitude and temperature envelope — engineered for the field, not the bench.

  • Open architecture · MOSASoftware + Compute

    Run the stack on FORGE or on your own NVIDIA hardware — and a trusted, NDAA-compliant supply chain behind all of it.

01 · Optics

The Optics: the Image the Sensor Receives

A sensor can only record the image the lens delivers to it. The optics decide how much of the sensor’s resolution is actually usable, whether that holds as conditions change in flight, and — through careful calibration — whether a pixel can be turned into a real-world direction.

At a Glance

  • Precision glass, not plastic, holds sharpness edge-to-edge and across temperature.
  • Resolving power is matched to the sensor, so no detail is lost at the lens.
  • Infrared-cut filtering keeps colour true for the model.
  • Metrically calibrated — every pixel maps to a known angle, the basis for georeferencing.

Glass versus plastic

Lens elements are made from optical glass or from polymer (plastic). Glass is optically superior and far more stable: it expands very little and holds focus as the air cools with altitude and as the camera heats and pressure changes. Plastic is lighter and cheaper, but it shifts with temperature — moving the focus and softening the image just as the environment changes.

Why it matters for AIA lens that holds focus across the whole flight envelope keeps every frame sharp. Consistent sharpness means consistent detections, rather than accuracy that drifts with temperature.
Sharpness across temperature sharpness cold hot glass — holds plastic — drifts

Resolving power

A 100 MP sensor only delivers 100 MP of real detail if the lens can draw detail that fine. Every lens has a resolving limit: below a certain size, fine features blur together. If the glass blurs beneath the pixel size, the extra pixels record blur rather than information. Matching the lens’s resolving power to the sensor grid is what makes a high pixel count real rather than nominal.

Why it matters for AISharp, fully-resolved pixels give a model true edges and texture. A soft lens feeds it approximations, and no amount of resolution downstream recovers detail the glass never formed.
Resolving fine detail Soft lens fine lines merge Matched lens lines stay distinct

Infrared-cut filtering

Silicon sensors respond to light well beyond human vision, into the near-infrared. Left unfiltered, that infrared light contaminates colour and softens focus, because it focuses at a slightly different point than visible light. An infrared-cut filter blocks it for accurate colour. Deliberately removing the filter — as on a monochrome path — opens the near-infrared back up for low-light and haze penetration.

Why it matters for AIClean, predictable spectral response means colour features a model relies on are trustworthy — and a dedicated near-infrared channel adds information exactly when visible light fails.
The spectrum a sensor can see visible near-infrared IR-cut filter blocks NIR → accurate colour Monochrome path removes the filter → opens NIR

Metric calibration

Every real lens distorts the image slightly. Metric — or photogrammetric — calibration measures that distortion and the exact internal geometry, so each pixel corresponds to a precisely known ray leaving the camera. Combined with the platform’s position and orientation, that is what turns a pixel into a coordinate on the ground: georeferencing, computed onboard rather than in a separate ground workflow.

Why it matters for AIIt changes the output from “there is a vehicle in the image” to “there is a vehicle at this location.” The calibration is the bridge from a pixel a model flags to a point on the map.
Pixel → ray → ground point lens ground (lat, lon)
How Sapient implements this

Aerial optics engineered to resolve the full 100 MP grid, athermalised to hold focus across the operating envelope, with controlled IR-cut response and metric calibration as the optical basis for onboard georeferencing. See the ROOK & ECHO sensors →

02 · Sensor

The Sensor: Turning Light Into Data

A digital image sensor is a grid of tiny light-collecting wells — the pixels. How large those wells are, how they are read out, and which light they respond to set a hard ceiling on image quality. Detail that is never captured cleanly cannot be recovered later, so the sensor also sets a ceiling on what an AI model can detect.

Sensor formats, drawn to scale Sapient 10K · 27.5 × 27.5 mm (square) Sapient · 10K 27.5 × 27.5 mm · square 1-inch type 13.2 × 8.8 mm · higher-end video 1/2.3-inch type 6.2 × 4.6 mm · typical drone cam LIGHT-COLLECTING AREA Large-format square sensor ≈ 6.5× a 1-inch sensor ≈ 27× a 1/2.3-inch drone cam Comparable area to full-frame — but square.
Sensor formats shown to scale. Sapient’s 27.5 × 27.5 mm square 10K sensor is a large-format sensor — comparable in area to full-frame, but square, and far larger than the 1-inch and 1/2.3-inch formats used in most drone cameras.

At a Glance

  • A large-format sensor gathers far more light, and full 12-bit RAW is kept end-to-end — so detail survives in deep shadow and bright highlight where 8-bit pipelines clip.
  • An electronic global shutter freezes fast motion with no rolling-shutter skew and no moving parts to wear.
  • Passive imaging: the sensor emits nothing, so nothing gives the platform away.
  • Resolution and sensitivity are balanced for wide-area detection, not just a pretty picture.

Sensor size and pixel size

A “full-frame” sensor measures 36 × 24 mm — many times the physical area of the small sensors in phones, action cameras and most video cameras. A larger sensor allows larger pixels, and a larger pixel is a bigger bucket: it gathers more photons before it fills up. More photons per pixel means a stronger signal relative to noise and a wider dynamic range — more usable detail in deep shadow and bright highlight at the same time.

Why it matters for AITargets hidden in shade or glare still carry the texture a model needs to detect and classify them. Detail lost to noise or blown-out highlights cannot be restored downstream.
Same light, two pixel sizes Full-frame pixel wide dynamic range Small pixel clips + noisier

Resolution: more pixels, or bigger pixels

Resolution is a trade-off. On a sensor of fixed size, more megapixels means smaller pixels — each collecting less light. Keeping pixels large means fewer of them. A physically larger sensor is what lets you have both: a high pixel count with pixels that are still large. What matters operationally is how many pixels land on the object — its ground sample distance. Too few and a target is an unresolvable blob; enough and it becomes a recognisable shape.

Why it matters for AIEvery model needs a minimum number of pixels on an object to classify it. Resolving fine detail across a wide area keeps small, distant targets above that threshold.
Pixels on target few pixels → blob many pixels → vehicle

Electronic global shutter

A sensor can be read two ways. A rolling shutter scans line by line, top to bottom, so on a moving platform straight edges skew and fast objects smear. A global shutter exposes every pixel at the same instant, recording one geometrically consistent frame. “Electronic” means it does this with no moving mechanical shutter — nothing to wear or fail — and with very short, precise exposures that freeze motion and reduce blur.

Why it matters for AIUndistorted, blur-free frames give detectors sharp, true edges — and a geometrically faithful frame is what makes it possible to place a detection accurately on the map.
Capturing a moving object motion → Rolling shutter skews & smears Global shutter one instant, true shape

Passive imaging

A camera is a passive sensor: it only receives ambient light. Active sensors — radar (including SAR) or lidar — must emit their own signal to see. That emission is itself detectable and can be jammed. A passive sensor transmits nothing to intercept, though it depends on there being light in the scene to begin with.

Why it mattersWhat the sensor emits is a survivability question, not an image-quality one — but it shapes where and how a passive imaging sensor can be used.
Receiving vs emitting Passive receives · emits nothing Active emits · detectable

Colour versus sensitivity

A colour sensor places a red, green or blue filter over each pixel (a Bayer array), so each pixel sees only part of the spectrum and full colour is reconstructed by interpolation. A monochrome sensor has no colour filters, so every pixel collects the full available light — giving more sensitivity, higher effective resolution, and reach into the near-infrared, which helps in low light and haze. The trade is colour information against raw light-gathering.

Why it matters for AIColour can be a useful discriminating feature; a monochrome or near-infrared channel can surface targets in conditions where a colour sensor sees almost nothing.
Colour filter vs no filter Colour (RGB) less light per pixel Monochrome more light + near-IR
How Sapient implements this

A 100 MP full-frame sensor (10,000 × 10,000 px) in RGB and monochrome variants — the monochrome path extending into near-infrared — with an electronic global shutter, operating fully passively. See the ROOK & ECHO sensors →

03 · Compute

The Compute: Why the AI Runs on the Aircraft

Where the AI runs is itself an enabling choice. Running it on the aircraft — at the edge — rather than sending imagery to a ground station or the cloud changes what is possible, and it is only possible because the compute is engineered to fit a flying platform.

At a Glance

  • The full 10K stream reaches the processor over a wide internal PCIe bus — nothing dropped or compressed.
  • Every pixel is processed onboard; only findings leave the aircraft, never raw video.
  • Decisions happen in the moment — no ground round-trip.
  • Serious AI in a sensor’s power budget — and it keeps working through a link blackout.

Moving the data: a high-bandwidth internal bus

A 10K sensor produces an enormous stream of data every second — far more than an ordinary camera connection can carry. PCI Express (PCIe) is a very wide, high-speed data path used inside computers; here it links the sensor directly to the processing module so the entire full-resolution stream reaches the GPU without a bottleneck. It is a separate internal path — not the radio link to the ground — and it is what makes onboard, full-frame processing possible in the first place.

Why it matters for AIThe model can run on every pixel of every frame, because the full stream actually reaches the processor — nothing is dropped or compressed on the way in.
Sensor to processor, full stream Sensor 100 MP · 10K PCIe · MANY LANES full 10K stream, uncompressed Processing module · GPU

Process onboard, not on the ground

A 10K video stream is far too much data to send over a radio link. Transmitting raw imagery is slow, floods the link, and stops entirely if the link drops. Processing on the aircraft means only the findings — a handful of small messages — ever need to be sent.

Why it matters for AIThe model runs where the data already is, so detection happens in real time and does not depend on a link back to base.
What crosses the link onboard compute link raw 10K too much findings ground

Answers in the moment

Sending imagery to the ground, waiting for it to be processed, and getting a result back all takes time. Deciding onboard removes that round trip — the answer is available while it still matters, not seconds or minutes later.

Why it matters for AIClosing the loop between seeing and deciding onboard is what makes real-time cueing and target tracking possible.
Time from seeing to deciding onboard see → decide ground see → send → process → return

A data-centre’s job in a sensor’s power budget

An aircraft offers very little power and no room for fans. Edge compute has to deliver serious AI performance within a few watts, in a small sealed package that sheds heat passively — and keep doing it as the camera warms up. Size, weight and power (together, SWaP) set the ceiling on how much AI can actually fly.

Why it matters for AIThe power and thermal budget decides how large a model, and how much resolution, can run in flight — not what a lab bench allows.
Performance within a flight budget GPU few watts · passively cooled full-frame AI in real time

Keep everything; keep working

Fast onboard storage retains every raw frame for detailed review after the flight, while only findings are sent live. And because the processing is self-contained, the mission keeps working even if the communications link is jammed or lost — findings simply queue until it returns.

Why it matters for AINothing is discarded for bandwidth, and the AI keeps running through a communications blackout rather than going dark with the link.
Retain everything, survive link loss raw frames retained link jammed / lost findings queue & continue

Many sensors, one clock

When more than one sensor flies on the same platform — several heads for a wider swath, or a mix of sensors — their frames only combine cleanly if they are captured at the same instant. Sapient synchronises capture across sensors to a shared timebase, so every frame shares a common trigger and timestamp rather than drifting apart from one exposure to the next.

Why it matters for AIFrame-accurate synchronisation lets the model fuse several sensors into one scene — stitching a wider area, comparing views, and holding a track as a target crosses from one sensor’s field into the next.
One trigger, frames aligned shared clock sensor 1 sensor 2 sensor 3 same instant, every sensor
How Sapient implements this

NVIDIA Jetson Orin NX (16 GB) edge compute on a Sapient-designed carrier board — with a 2 TB NVMe SSD for full-mission RAW retention, flight interfaces, and a sealed 160 g package (190 × 55 × 30 mm) that mounts anywhere. See the FORGE module →

04 · Software

The Software: From Frame to Finding

Hardware captures the image; software turns it into an answer. The enabling ideas here are less visible than glass and silicon, but they decide whether an AI model can run on the aircraft at all, how well it performs, and whether its output is usable the moment it lands.

At a Glance

  • Your model runs unchanged inside an isolated, encrypted container.
  • Inputs are shaped to exactly what the model expects — scaled, tiled, calibrated.
  • Fast enough to keep pace with the sensor in real time.
  • Outputs speak the receiver’s language — Cursor-on-Target (CoT), straight into ATAK and C2.

Containers: a sealed box for the model

A container packages an application together with everything it needs to run, so it behaves identically wherever it runs and cannot reach anything outside its box — like a shipping container: standard on the outside, your cargo sealed inside. Each AI model runs in its own container. You bring a model as-is, updating it means swapping the box rather than rebuilding the system, and because the box is sealed the model and its data stay isolated.

Why it matters for AIA model trained anywhere runs unchanged — no porting, no retraining — while its weights and your data stay private inside the container.
Your model, in a sealed container HOST · FORGE Your AI model isolated & sealed update = swap box

Feeding the model what it expects

An AI model is trained on images of a fixed size and a particular scale — how much ground each pixel covers. A 10K frame is far larger than that, and its scale changes with altitude. The software cuts each frame into tiles the size the model was trained on, and resamples them to the scale it learned, so every tile looks familiar. This is possible because the optics are calibrated and each frame carries its own pose, so the ground scale of every pixel is known.

Why it matters for AIDetectors are far more accurate on input that matches their training size and scale. Matching it in software lets a model perform well on the sensor without being retrained for it.
Tile & rescale to the model 10K frame → tiles rescaled tile model

Fast enough for real time

Running a large model on every frame, in the air, on a few watts, means tuning the model to the exact onboard processor — using lower-precision arithmetic and fused operations that do the same work in far fewer steps. This is the difference between analysing a live stream and steadily falling behind it.

Why it matters for AIAcceleration is what lets a full-resolution model keep pace with the sensor in real time, instead of dropping frames or shrinking the image to cope.
Frames processed per second as-is accelerated same processor · same power budget

Speaking the receiver’s language

A finding is only useful if the systems downstream can read it. Rather than inventing a new format, the software emits findings in the open message standards and file types command systems already use — so nothing new has to be installed on the ground to receive them.

Why it matters for AIThe model’s output arrives as a map-ready track in the operator’s existing tools, not as a proprietary file someone has to translate first.
Findings in open standards finding CoT JSON existing C2 / ATAK
How Sapient implements this

Models run in isolated Docker containers, accelerated for the edge with TensorRT, with native-resolution crops kept in the open RAW DNG format, and findings delivered as Cursor-on-Target (CoT) into ATAK and existing C2. See the IGNITE:AI framework →

Designed together

Capabilities No Single Layer Could Deliver

The defining capabilities of the platform sit between the layers — they exist because optics, sensor, compute and software are engineered as one system.

CALIBRATION02 · OPTICS POSE · IMU + GNSS04 · COMPUTE 10K FRAME01 · SENSOR ADAPTIVE INPUT GSD KNOWN PER PIXEL GPU-ACCELERATED MODEL-READY TILES TILE SIZE · OVERLAP · GSD MATCHED TO THE MODEL
Adaptive input: because the optics are calibrated and every frame carries its pose, ground sample distance is known per pixel — so the 100 MP frame is retiled and resampled to each model’s trained tile size, overlap and GSD. No retraining to onboard a model. Pre-processing is GPU-accelerated.

See Why the Engineering Matters

Explore sets out what these technologies buy you in the field — coverage, altitude, and a near-silent RF footprint.