Back to Boğaziçi

04Boğaziçi Savunma Teknolojileri · Long-term internship

More thana box.

Finding a fixed-wing UAV in the frame was the first step. Giving it a lasting identity, drawing its trail, logging every frame and estimating its distance came next.

Model
YOLO26L · 1280 px fine-tuning
Tracking
ByteTrack · persistent IDs
After fine-tuning
P 0.97 · R 0.96 · mAP@50 0.975
Speed in the recording
≈25 FPS · 1080p video

My contribution

I collected data in the field to the project’s requirements and turned the videos into a labelled dataset. I trained YOLO26L and fine-tuned it on new field data, and built the ByteTrack tracking, the detection console, the distance-estimation module and the interface. So that confidential data never went to the cloud, I wrote a local annotation tool.

04Real-time trackingYOLO26L · ByteTrack
A 20-second recording of the application, looping: two UAVs as #1 and #2, detection on a grass strip, tracking with the console over hills, a distant target over a lake. The speed and altitude gauges belong to the source video. Open-source flight footage.

01Data

From the field to a dataset.

Fixed-wing platforms look different from multirotors: long wings, low flight, high speed. The drone model would not be enough for them; data from real flights was needed.

I collected real flight videos in the field to the project’s requirements and turned them into a dataset. The first set came from 11 videos: daylight and dusk; countryside, coast and farmland; different flight angles. Consecutive frames of a 60 FPS recording are nearly identical, so I kept every 20th frame: about 4,800 frames.

The data was confidential and could not be uploaded to the cloud, so I labelled the frames with my own fully local annotation tool: the model proposes a box, and I confirm, correct or delete it. A second round added about 7,200 frames from 11 new videos. That tool later became LabelMate.

Explore the LabelMate case

01Frame samplingExplanatory
Explanatory drawing: keeping every 20th frame of a 60 FPS recording (stride = 20).

Confidentiality

The field data used for training is confidential under contract. Every image on this page comes from open-source flight videos I played through the application to show how the system works.

02Evaluation

Split videos, not frames.

If consecutive frames of one video sit in both training and testing, the model memorises the scene instead of generalising.

In the new dataset I kept all frames of each video in a single split: 7 videos for training, 2 for validation and 2 for testing. I fine-tuned from the previous training’s weights with a 1280-pixel input: a low learning rate (0.0004), reduced mosaic (0.10), mixup and copy-paste off, and part of the old data kept in training to avoid forgetting.

02Video-level split7 / 2 / 2
Explanatory drawing: each video sits wholly in one split (report, Table 4.10).
02Fine-tuningYOLO26L
  • First training · 640 px
  • Fine-tuned · 1280 px
Approximate values from the report (Tables 4.8 and 4.12). The two trainings differ in data and input size; the change cannot be attributed to fine-tuning alone.
MetricFirstFine-tunedChange
Precision≈0.96≈0.97+0.01
Recall≈0.86≈0.96+0.10
mAP@50≈0.91≈0.975+0.065
mAP@50-95≈0.66≈0.77+0.11
Approximate values from the report (Tables 4.8 and 4.12). The two trainings differ in data and input size; the change cannot be attributed to fine-tuning alone.Recall: how many real UAVs were caught. Precision: how many detections really were UAVs.

03Tracking

An identity for every target.

Detecting each frame on its own cannot say which UAV is which when two appear at once.

03Tracking recording01 / 04
Tracking off: a fixed-wing UAV close to the camera on a grass strip, in a green UAV box with no identity.

With tracking off, detection only: a box, but no identity.

Two UAVs over wooded land seen from the air, in an orange #1 box and a blue #2 box; the detection console at the top right.

Two UAVs tracked as #1 and #2. The speed and altitude gauge belongs to the source video.

A UAV with identity #1 in an orange box over green hills; at the top right a console lists position, size and distance for every frame.

For every frame the console lists the position, the box size and the estimated distance.

A small UAV marked #1 over a misty lake and hills; the console gives the box as 43 × 7 pixels.

The identity holds on a target a few pixels wide; the console gives the box as 43 × 7 pixels.

With tracking off, detection only: a box, but no identity.

I added ByteTrack to the pipeline: detections are linked to the previous frame’s tracks with a Kalman filter and IoU matching, and low-confidence detections stay in the tracking. Each target’s past centres are drawn as a thin trail, so the direction of flight reads at a glance.

Track memory, matching threshold, the lower confidence bound and trail length are set in the interface (230 frames, 0.90, 0.35 and 40 frames in the recording). In the recordings the interface processes about 25 frames per second (1080p video).

Tracking quality was not evaluated with a measure such as MOTA, HOTA or identity switches; the recordings show the behaviour.

04Observability

A numerical record of what is seen.

Looking at boxes was not enough for debugging. I added a Detection Console that writes out every inference frame line by line. False detections, the continuity of identities and the distance estimate can all be checked in those lines.

04Detection consoleframe 23289
The Detection Console window: green text with frame numbers, times and #1 UAV lines; each line gives the confidence, box corners, centre, size and approximate distance.
The console window, taken from the recording (12.5 s).

One line, field by field

F23289frame388.15stime in the video#1track IDUAVclassconf=0.92confidencex1=513 y1=512 x2=733 y2=550box cornerscx=623 cy=531centre220x38size (px)~5mestimated distance

05Decision support

Approximate distance from a box.

No LiDAR, stereo camera or radar: one camera and the detection box.

In the pinhole camera model an object’s width in the image is inversely proportional to its distance: d = W · f / w. W is the target’s real wingspan, f the focal length in pixels and w the width of the box.

I put three methods into the interface: the pinhole calculation from wingspan and focal length; a single K constant from a box measured at a known distance (d = K / w); and, with no calibration at all, a Far / Medium / Near indicator from the box’s share of the frame. The result appears in the console and in the tracking panel.

The interface’s settings rows: Bellek 230 frames, Eşleşme 0.90, Conf lower bound 0.35, İz 40 frames; Mesafe 1 · Pinhole, Kanat 1.8 m, f 550 px; the processing rate of 25.6 fps below.
The settings in the recording: distance = pinhole, wingspan = 1.8 m, f = 550 px.
05Pinhole calculationW 1.8 m · f 550 px

Estimated distance≈ 4.5 mConsole ~5m

The 220 × 38 box from the console: 1.8 × 550 / 220 ≈ 4.5 m (the console writes ~5m).

With no calibration: the box’s share of the frame

  1. Far< 0.08%
  2. Medium0.08% – 0.40%
  3. Near≥ 0.40%
Interactive explanation: the same calculation as the recording’s settings. The real wingspan of the aircraft and the focal length of the camera in this open-source video are unknown, so the number shows the method, not a real distance.

Its limits

  1. 01

    It assumes the target is roughly perpendicular to the camera; in an angled pass the box narrows and the distance comes out too large.

  2. 02

    Lens distortion and perspective are not modelled.

  3. 03

    The result depends directly on calibration: an error in f or K passes straight into the distance.

  4. 04

    For very distant targets, an error of a few pixels becomes a large relative error.

The module is not a measuring sensor but an aid that gives the operator context. It was not compared with real distances in the field.

06The outcome

From detection to decision support.

  1. 01

    Field data

    A dataset collected to requirements and split by video.

  2. 02

    Model

    YOLO26L, fine-tuned at 1280 px on new field data: P ≈0.97, R ≈0.96, mAP@50 ≈0.975.

  3. 03

    Tracking and console

    ByteTrack identities, motion trails and a per-frame detection log.

  4. 04

    Distance

    Pinhole, a K constant and relative proximity: three methods, one interface.

A desktop prototype. Porting to Jetson Orin NX and INT8 are future work. Tracking quality and distance accuracy were not measured numerically.

Contact

Let’s buildthe next system together.

As an engineer focused on turning AI models into systems that work in the real world, I’m open to new opportunities and technical collaborations. If you work on computer vision, edge AI or applied machine learning, I’d be glad to connect.

Muhammed Ali Yıldırım

Applied AI / ML Engineering

Back to top