Back to the work

01Edge AI · Thesis · Two-person team

AI-poweredsearch & rescuesystem

From a camera in the air to the operator’s screen: a working human-detection system, tested in the field.

On-device inference
Jetson Nano · TensorRT FP16
Selected model (validation)
YOLOv11 · mAP@0.5 0.8694
Stream to operator
DeepStream · RTSP
Field test
Gazi University campus

My contribution

The system was built end to end by a team of two—from research and procurement to sponsorship and assembly. I was primarily responsible for the computer vision, dataset and model development, training and evaluation, camera selection and the Jetson-based edge AI pipeline. The project was supported by Gazi University’s Scientific Research Projects unit (BAP) and has been completed.

01Field test · real flightIMX477 · 160°
Field test, Gazi University campus. Real flight footage from the drone camera; the boxes and confidence scores are the model’s output.

01Problem

Finding a person from the air.

In search and rescue, time is critical, and a drone can cover a wide area quickly. But from altitude a person becomes a small part of the frame. Altitude, background and visibility keep changing; a person may be partly hidden or at the edge of the view.

The system does not decide. It marks candidate people on the image; confirmation belongs to the operator. The goal is to reduce the load of scanning video alone.

01Field test · framehuman 0.63
Wide-angle aerial view of a football pitch; the model has marked two people who appear very small in the frame as human.
A frame from the field test. The 160° wide-angle lens covers the pitch in a single frame; the person is a small part of it.

02System

Flight and perception, deliberately apart.

Two paths, one ground station.

02System architectureRedrawn
Redrawn from the project’s system architecture diagram.Open the original diagramproject image · English
02Platform · in the fieldHolybro X500 V2
A four-rotor drone carrying electronic components, standing on the grass of a football pitch.
The system’s UAV platform at the field test.

Flight is flown from an RC transmitter through a Pixhawk 6C flight controller; safety functions such as return-to-home (RTH) and automatic take-off make it semi-autonomous. Image processing runs on a separate unit, a Jetson Nano. There is no command or control signal between the two: the AI only produces information and supports the operator’s decision.

03Model

Same data, two models.

Which model should run on the Jetson Nano?

On the PROJE dataset, compiled from the SARD and WiSARD datasets, YOLOv8 and YOLOv11 were compared under the same data splits and similar training procedures: pretrained weights, 640×640 input, 100 epochs, Google Colab Pro.

A data decision: informativeness over volume

WiSARD frames come at resolutions up to 4096×2160, but training runs at 640×640. Frames in which a person would shrink below what the detector can learn after that resize were deliberately excluded during curation.

03YOLOv8 · YOLOv11Validation
  • YOLOv8
  • YOLOv11 · selected
PROJE dataset · end-of-training validation metrics, both models with augmentation (thesis, Table 5.1). Scale 0–1.
MetricYOLOv8YOLOv11Change
Precision0.89970.9109+0.0112
Recall0.79530.8065+0.0112
mAP@0.50.84330.8694+0.0261
mAP@0.5:0.950.40130.4481+0.0468
PROJE dataset · end-of-training validation metrics, both models with augmentation (thesis, Table 5.1). Scale 0–1.Precision: how many marked boxes are really people · Recall: how many people were found · mAP@0.5:0.95: average quality including tighter box placement

Detections on dataset images

Examples from both sources of the PROJE dataset: targets at the frame edge, partly visible, or very small.

03Dataset samplesSARD · WiSARD
  1. Aerial view of a grassy bank and shrubs; a partly visible person at the left edge is marked as human with 0.78 confidence.

    SARDA partly visible person among vegetation at the left edge of the frame (0.78).

  2. Top-down view of three cars parked on grass; a partly visible person at the right edge is marked as human with 0.61 confidence.

    SARDA partly visible person at the right edge, next to parked cars (0.61).

  3. Oblique aerial view of a ploughed field and a tree line; three small people at the foot of the trees are marked as human.

    WiSARDThree people at the foot of a tree line, each a tiny part of the frame; all three are marked.

Boxes and confidence scores are the model’s output. Images: the SARD (Sambolek and Ivasic-Kos, IEEE DataPort) and WiSARD (Broyles, Hayner and Leung, IROS 2022) datasets; rights remain with their owners.

Why YOLOv11?

  1. 01

    Better-placed boxes

    The largest gain is in mAP@0.5:0.95, which covers stricter IoU thresholds: 0.4013 to 0.4481. Boxes sit more accurately on the person.

  2. 02

    Fewer false alarms

    In the confusion matrix, background mistaken for a person fell from 821 to 631.

  3. 03

    Misses in a similar range

    Missed person instances stayed close, at 1,487 and 1,562. The choice rests on precision, mAP and box placement.

YOLOv11 (with augmentation) was selected for deployment on the Jetson Nano.

04Pipeline

From frame to operator.

The model is prepared once and moved to the device; every frame is processed on the device and streamed to the operator.

04DeepStream pipelineJetson Nano

Preparation · once

  1. TrainingPyTorch · Ultralytics · Google Colab Pro
  2. ONNXportable model format
  3. TensorRT FP16device-specific inference engine

Runtime · on the Jetson Nano, for every frame

  1. CameranvarguscamerasrcIMX477 · CSI
  2. Batchingnvstreammuxframes grouped for inference
  3. InferencenvinferYOLOv11 · TensorRT FP16
  4. Boxes + scoresnvdsosddrawn onto the video
  5. Encodingnvv4l2h264enchardware H.264
  6. Streamingrtph264pay · RTSPover the network
  7. Ground stationRTSP client (e.g. VLC) · operator confirmation
Plugin names come from the project’s own DeepStream diagram, which describes itself as illustrative; the runtime configuration was not published. Format conversions (nvvideoconvert) are left out of the drawing. No on-device FPS or end-to-end latency measurement was reported.Open the original diagramproject image · English

05Field

Real flight, real footage.

Do the results hold up in the field?

05Field test01 / 03
Wide-angle aerial view of a football pitch; two people, one at the bottom edge of the frame, marked as human.

Medium altitude, wide angle: two people, one at the bottom edge of the frame (0.63 · 0.69).

Aerial view of the pitch from another angle; two people, one far away and one at the bottom edge, marked as human.

A different position and distance: two people (0.50 · 0.72).

Aerial view of the pitch; one person in the centre marked with high confidence, one partly visible at the lower-left edge marked with low confidence.

A partly visible person at the frame edge keeps a low confidence score (0.79 · 0.29).

Medium altitude, wide angle: two people, one at the bottom edge of the frame (0.63 · 0.69).

The first field tests took place on the football pitch of Gazi University’s campus, with the required permissions. The wide, open area allowed different altitude and distance combinations to be tried safely.

Although the target’s size in the frame changed markedly, the model detected the human class consistently, including people at the edge of the frame and partly visible.

Field validation is qualitative: the footage was assessed by direct observation. No field recall, FPS or latency value was reported.

06Outcome

More than a model: a working system.

  1. 01

    From data to model

    Data compiled from two public datasets; two architectures compared under the same conditions, and a reasoned model choice.

  2. 02

    From model to device

    Inference on a Jetson Nano with ONNX and TensorRT FP16, inside a DeepStream pipeline.

  3. 03

    From device to operator

    Annotated video streamed to the ground station over RTSP; a perception path kept separate from flight.

  4. 04

    From the lab to the field

    Qualitative field validation on campus with real flight footage.

Bachelor’s thesis · Gazi University Faculty of Technology · supported by Gazi University BAP · 2026 · joint work with Mohamedou Mohamedhen Vall

Contact

Let’s buildthe next system together.

As an engineer focused on turning AI models into systems that work in the real world, I’m open to new opportunities and technical collaborations. If you work on computer vision, edge AI or applied machine learning, I’d be glad to connect.

Muhammed Ali Yıldırım

Applied AI / ML Engineering

Back to top