Back to Boğaziçi

04Boğaziçi Savunma Teknolojileri · Long-term internship

Droneor bird?

A target a few pixels wide in the sky. First catch it, then decide what it is.

Architecture
Detection + classification
Models
YOLO26m · YOLO11m-cls
Classes
Drone · bird · background
Speed (RTX 3060)
12.8 → 26.0 FPS

My contribution

I merged and checked the open-source datasets, trained the detection and classification models, built the two-stage inference pipeline and the desktop interface, and accelerated the pipeline with TensorRT FP16.

04Drone / birdYOLO26m · YOLO11m-cls
The project’s demo clip (6 s): the drone the system marked in green, the bird in blue. Open-source footage.

01Data

Distant, small and alike.

From a distance, a drone and a bird are both a smudge of a few pixels. I reviewed open-source datasets for label integrity, empty and faulty labels, scenario fit and class balance, and merged them into one set.

The first model trained on it raised false alarms on small targets and busy backgrounds. The WOSDETC / Drone-vs-Bird Challenge dataset was then added under a data usage agreement. The agreement forbids sharing it, so it does not appear on this page.

I used CVAT and Roboflow for labelling. All data passed the same checks before training: image–label pairing, a visual review of the boxes on the images, removing broken files and converting to the YOLO format. Faulty labels went to zero.

01Sample outputdrone · bird
A cloudy sky over a field. Near the horizon on the lower left a drone is marked in a green box; on the right, above the trees, a bird in a blue box.
A drone (green) and a bird (blue) marked by the two-stage system in the same frame. The loupes magnify those two boxes. Open-source image.
The merged open-source dataset
SplitImagesBoxes
Train10,45015,517
Validation2,7494,082
Test1,3562,013
Total14,55521,612

10,159 of the boxes are drones and 11,453 birds.

02Architecture

Catch first, then decide.

When one model has to both find and tell apart, small targets slip away or land in the wrong class. The job was split in two.

02Two-stage pipelineRedrawn

Once: TensorRT FP16 engines

  1. Video detectionbest_target_640.engine640 × 640 input
  2. Image detectionbest_target_960.engine960 × 960 input
  3. Classifierbest_wbg_224.engine224 × 224 input

Every frame

  1. ReadreadOne frame from the video
  2. Detectdet · YOLO26mOne class: target. Catch everything first.
  3. CropprepOne crop per box
  4. Classifycls · YOLO11m-clsDrone, bird or background
  5. DrawrenderBox and class label
  6. ShowdisplayLive view in the interface
  7. WritewriteRecording the output
The stage names are the ones the interface times on every frame (read, det, prep, cls, render, display, write). Redrawn.Open the project’s original diagrampipeline_overview.png · 1600 × 650

The detector puts drones and birds into a single “target” class and concentrates on catching everything that flies. The decision is made in the second stage, by a classifier that sees only that box’s crop: the detector’s high recall and the classifier’s precision meet in one pipeline.

03Error analysis

A third answer: neither.

The first classifier knew only drones and birds. When the detector mistook a patch of sky, the edge of a leaf or a shadow for a target, that crop had no choice but to land in one of the two classes.

I reviewed the detector’s false positives and added them to the classification data as a “background” class. The three-class model can reject those crops; only drones and birds reach the interface.

The report states, from observation in tests, that the two-stage design clearly reduced false alarms, especially birds taken for drones. No separate figure was measured for that reduction.

03Decision tableExplanatory
What the crop showsTwo classesThree classes
Dronedronedrone
Birdbirdbird
Sky, leaf edge, shadow (false detection)drone or birdbackground → rejected
Explanatory table: how the same three crops pass through a two-class and a three-class classifier.

04Interface

The model, inside an application.

04Desktop interfacePyTorch FP16
The Video Test – 2-Step (GPU + Batch CLS) window. A drone low over buildings is marked in a green box; the processing rate and stage times are written above.
A screenshot of the interface in PyTorch mode (GPU + Batch CLS), frame 178. Open-source test video.

I built a desktop interface that joins the two models in a single flow. An image or a video is loaded; on every frame detection runs first, then classification, and the result appears with its box and class. The interface also shows the live processing rate and the time of each stage.

The window’s top lines

İşlem12.4 fpsHedef25.0 fpsDrone1Bird0crops1frame178

read3.8det27.3prep0.1cls6.1render5.3display2.5write10.1total58.2

Instant values for a single frame. det and cls are the two models, total is the frame’s whole time (ms).

İşlem
Frames processed per second
Hedef
The real-time target: 25 FPS
crops
Crops sent to the classifier

05Optimisation

The target was 25 FPS.

With PyTorch FP16 the pipeline processed 12.8 frames per second: about half the real-time target.

PyTorch FP16
12.8FPS
TensorRT FP16
26.0FPS

I converted both models to TensorRT FP16 engines; an engine is built once and then runs directly on the GPU. The pipeline reached 26.0 FPS. The largest gain was in classification: with 224 × 224 crops processed in a batch, its time fell from 15.4 ms to 2.6 ms.

05Stage timesms / frame

PyTorch FP1687.0 ms

TensorRT FP1644.8 ms

Same test video and hardware: RTX 3060, CUDA 12.6 (report, Table 4.5). “Other” is the part of the total the report does not itemise, computed as the total minus the listed stages.
MeasurePyTorch FP16TensorRT FP16Change
Frames per second12.8 FPS26.0 FPS×2.03
Total latency87.0 ms44.8 ms−48.5%
Detection25.2 ms18.9 ms−25.0%
Classification15.4 ms2.6 ms−83.1%
Render6.4 ms4.4 ms−31.3%
Display3.0 ms1.9 ms−36.7%
Write15.9 ms12.5 ms−21.4%
Same test video and hardware: RTX 3060, CUDA 12.6 (report, Table 4.5). “Other” is the part of the total the report does not itemise, computed as the total minus the listed stages.Open the project’s original charttensorrt_comparison.png · 1600 × 720

A desktop measurement. Accuracy was not compared separately between the two modes. Jetson Orin NX was researched as the target platform; embedded deployment and INT8 are future work.

06The outcome

From a model to a working pipeline.

  1. 01

    Data

    Merged and checked open-source datasets, brought into one YOLO structure.

  2. 02

    Two stages

    High-recall detection, classification per crop.

  3. 03

    Background class

    A third class learned from the detector’s false positives.

  4. 04

    TensorRT FP16

    From 12.8 to 26.0 FPS, with stage times measured live in the interface.

A desktop prototype (RTX 3060). The public repository contains the shareable two-stage subsystem; WOSDETC data and company data are not shared.

Contact

Let’s buildthe next system together.

As an engineer focused on turning AI models into systems that work in the real world, I’m open to new opportunities and technical collaborations. If you work on computer vision, edge AI or applied machine learning, I’d be glad to connect.

Muhammed Ali Yıldırım

Applied AI / ML Engineering

Back to top