Investigating how to help an autonomous bicycle reliably stay in its lane using camera-based perception and real-time processing on an NVIDIA Jetson Orin Nano Super.
Aman NindraRayan
University of California, Merced
Aman: Computer vision, dataset preparation, model training, inference optimization, and perception integration.Rayan: Bicycle robotics and control hardware, including the ESP32-S3 controller, steering feedback, and 250 W motor system.
The UC Merced Honors poster from an earlier stage of the project. Its model and system descriptions predate later experiments documented below.
The research question
How can I build a camera-based perception system that reliably identifies the bicycle’s lane and provides a useful steering target within the Jetson’s real-time compute budget?
To investigate this problem, we tested several fundamentally different perception approaches, ranging from semantic segmentation and multi-task perception networks to anchor-based lane detection.
Some models performed extremely well on benchmark and offline test data, but failed to reproduce those results on footage captured from the bicycle. This turned what initially looked like a lane-detection problem into a much larger set of research questions spanning perception, geometry, domain shift, embedded inference, and control.
Questions we had to answer
01
Is there enough compute on the Jetson Orin Nano Super to run the perception system in real time?
02
How do we determine which detected boundaries belong to the bicycle's current lane?
03
What should the system do when only one lane boundary, or no lane markings at all, are visible?
04
How can road segmentation provide guidance when explicit lane detection fails?
05
Does the difference in camera height and viewpoint between the training dataset and bicycle camera create a significant domain gap?
06
What data augmentation best reproduces real conditions such as glare, shadows, motion blur, exposure changes, and camera vibration?
07
How can the model be optimized with ONNX, TensorRT, and FP16 without changing its predictions?
08
How do we convert perception output into stable lane geometry and steering commands while keeping predictions synchronized with the correct camera frame?
These experiments exposed an important distinction between offline accuracy and real-world reliability. A model could produce convincing results on a test dataset while still failing once mounted on the bicycle.
The failure could occur at several different stages of the pipeline: the model might not recognize the markings, postprocessing might select the wrong boundaries, useful predictions might disappear when only one boundary is visible, frames might become stale before reaching the controller, or coordinates might become inconsistent between the model input, camera frame, and steering system.
This meant that improving benchmark accuracy alone was not enough. The actual problem was determining how perception, postprocessing, temporal consistency, embedded inference, and control could work together reliably enough to keep the bicycle on the road.
Where the project stands?
For lane detection, our current model produces a high F1 score (0.79) on the test dataset, but real-world results are less reliable when the camera jitters, its height differs from the training viewpoint, or lane markings are partially faded.
In this LaneATT example, the model working really good because our model is trained with Data Augmentatios that are suited for low-light data. In addition, the lanes are clearly visible.LaneATT on highway videoLaneATT on highway videoLaneATT on UC Merced roads
Ego lane selection
Ego lane selection
Implemented, not formally evaluated
Given several candidate lane boundaries from the detector, this step decides which two belong to the bicycle's own lane and builds a midline between them.
Why I tried it. A lane model outputs a set of boundary curves per frame, not a labeled left/right pair. Something has to turn that set into the one corridor the bicycle should follow before it can be used for steering.
What I implemented.
Restrict every candidate lane to the vertical range shared by all of them, so curves are only compared where they overlap.
Resample each candidate onto the same 100 y-positions with a second-order polynomial fit (x as a function of y), so unevenly sampled curves become directly comparable.
Measure each resampled lane's average horizontal offset from the image's center column, both signed and absolute.
Split candidates into left and right groups by the sign of that offset, then keep only the closest candidate on each side.
Average the surviving left and right x-positions at each shared y-value to synthesize a midline between them.
What I learned?. The module keeps two functions, get_ego_lanes2 and get_ego_lanes3, that are exact duplicates of each other; only one is actually needed.
Configurations, files, and additional videos
How selection can fail
Fewer than two candidate lanes, or any candidate with fewer than three points, returns no result.
If the candidates' vertical ranges don't overlap at all, there is no shared region to compare and the function returns no result.
If every remaining candidate lands on the same side of center, there is no left/right pair to choose from and the function returns no result.
Files
LaneATT/lib/ego_lanes.py
Experiments and approaches
Each entry follows the same structure: why I tried it, what I implemented, the evidence, what went wrong and why, and what I learned. Status labels are specific: evaluated means measured on a dataset or checked offline; partially implemented means code exists but training, validation, or integration is incomplete; exploratory means investigation without a completed result. None of these is a claim that the approach cannot work.
Lane detection
The core question: can a model recover the bicycle’s lane boundaries from its own camera? The approaches run roughly in the order I tried them.
LaneNet and H-Net
Evaluated offline
Segment lane pixels, cluster them into lanes, fit curves, then pick the ego lane.
Why I tried it. LaneNet splits the problem into which pixels are lane markings (binary segmentation) and which lane each pixel belongs to (instance embeddings), so it handles a varying number of lanes. H-Net learns a transformation that makes lane points easier to fit with a polynomial.
What I implemented.
ENet-based LaneNet with binary-segmentation thresholds, mean-shift-style clustering of embeddings, minimum cluster size and distance settings, and a cap on retained lanes.
Raw versus sigmoid-transformed embeddings.
Polynomial fitting, left/right ego-lane selection, temporal smoothing, and holding previous estimates with decay and jump rejection.
Separate H-Net training on TuSimple lane points with third-order fitting, and inference with and without H-Net.
Later CULane loaders: LaneNet (image, binary mask, instance mask rasterized from .lines.txt) and H-Net (lane-point supervision).
What went wrong?. I could not get stable, correctly assigned ego-lane boundaries. Cluster identities and fitted curves changed between frames, so a boundary could bend away or vanish while the marking was still visible. The notebook also records two implementation failures: checkpoint keys that did not match the instantiated architecture, and H-Net receiving unsigned-byte input where float tensors were required.
Why?. The pipeline chains segmentation, clustering, curve fitting, and temporal association; a change at any stage changes the final boundary. A lane-shaped cluster is not automatically the left or right ego boundary. Applying a sigmoid changes distances in the embedding space the instance loss was trained on, which can alter clustering. The checkpoint and dtype errors are compatibility problems, separate from predictive quality.
What I learned?. Every stage between the network and the boundary is another place for the answer to change. This pushed me toward a model that predicts lane geometry directly. A learned perspective transform still depends on correct preprocessing and reliable detected points; it does not fix camera-domain differences.
CULane masks are rasterized from .lines.txt at original resolution and are not necessarily identical to the distributed CULane segmentation PNGs.
The CULane loaders exist, but I have not trained and deployed a new CULane LaneNet/H-Net model with them.
LaneATT
Evaluated on CULane · main approach
An anchor-based model that predicts lane geometry directly. It became my main lane detector.
Recorded LaneATT inference on bicycle-camera video; this is a qualitative excerpt, not a coverage evaluation.
Why I tried it. LaneATT removes LaneNet’s clustering stage. It samples backbone features along predefined lane anchors, uses attention across anchors, and predicts lane confidence plus geometric corrections to each anchor, followed by lane-specific non-maximum suppression.
What I implemented.
CULane training through the shared loader (culane.py → lane_dataset_loader.py → lane_dataset.py), transforming the image and lane points together.
Standard configuration: 640 × 360 input, 72 vertical sample positions, 1,000 anchors, raw output of about (1, 1000, 77).
Backbone comparisons from ResNet18 to ResNet152. Comparable training metrics exist for ResNet18/34/50; the larger variants have saved inference outputs only.
Training on UC Merced’s Slurm GPUs (current scripts request L40S; saved benchmark notes include A100 runs) and AWS SageMaker experiments.
Offline tools for replaying video, comparing checkpoints, and frame-level debugging (inference.py, fastLane.py, lib/video.py, LaneATT_debug.ipynb).
Experiment
Best validation F1
End-of-run test F1
ResNet18 Aug2
0.7770
0.7549
ResNet34 Aug2
0.7852
0.7682
ResNet50 Aug2
0.7807
0.7402
ResNet34 (Sept 10)
0.7779
0.7430
CULane results. Best validation F1 and end-of-run test F1 are separate recorded evaluations; the test column is not necessarily the best-validation checkpoint. ResNet34 Aug2’s best validation checkpoint was epoch 13.
What went wrong?. Strong CULane numbers did not carry over to the bicycle. My field estimate is that LaneATT produced usable left and right boundaries in fewer than about 25% of bicycle-camera frames, even when markings were visible. That figure comes from reviewing footage, not from a labeled evaluation.
Why?. Camera-domain shift is my strongest hypothesis: lower mounting height, different pitch and field of view, vibration and roll, different surfaces and markings, shadows, glare, motion blur, and thin distant markings after resizing. It has not been isolated. Differences between preprocessing paths, ego-lane selection after inference, and stale visualization could also contribute (see failure analysis).
What I learned?. A 0.785 validation F1 on CULane says little about bicycle footage. A bigger backbone did not help either: ResNet50 did not surpass ResNet34 in my runs.
Configurations, files, and additional videos
What the training path taught me
A .lines.txt file is the training target, not a substitute for the image.
Only the configured annotation path is read; extra copies of annotation folders are ignored.
Geometric augmentation must move the lane points with the image.
Preprocessing must match the checkpoint. This pipeline keeps OpenCV BGR order rather than converting to RGB.
The ~0.79 F1 often quoted for this project is the best CULane validation result (ResNet34 Aug2), not bicycle accuracy.
Trying to make training imagery look more like what the bicycle camera sees.
Why I tried it. If the gap between car and bicycle cameras is the problem, augmentation is the cheapest lever to try before collecting new labels.
What I implemented.
Rotation, translation, scaling, random perspective, horizontal flips, brightness/contrast, gamma, shadows, sun flare, motion blur, Gaussian noise, compression, and coarse dropout.
What went wrong?. While, the LaneATT did perform better in the real-world test data, it isn't enough for reliable Lane Detection
Why?. The Model is too weak for real-world data even if you use specific data augmentatin techniques for video. Need a stronger model, not just retraining but with better data augmentation
What I learned?. Data Augmentation can help, but isn't usefull for overall effectiveness
Road and lane segmentation
A parallel direction: segment the drivable road and derive a travel corridor from it, which might work even when individual markings are unreliable.
HybridNets
Evaluated offline · field FPS reported
Joint road and lane segmentation, plus a geometric pipeline that turns the road mask into a steering target.
Why I tried it. HybridNets combines road-scene tasks in one network. A drivable-area mask could supply a path when individual markings are faint.
What I implemented.
Training, validation, video and camera inference, and deployment experiments. Some recorded runs froze the detection branch and trained segmentation only.
Record
Road IoU
Lane IoU
Lane precision
Lane recall
metrics.jsonl, epoch 19
0.8401
0.2552
0.2666
0.7001
meter2.jsonl, epoch 29
0.8388
0.2611
0.2715
0.7028
Saved validation records (background / road / lane classes). Road is far easier than thin lane markings; lane recall is high but precision is low, meaning many false-positive lane pixels.
What went wrong?. Lane quality was much weaker than road quality. In my field measurement HybridNets ran at about 12 FPS after TensorRT optimization, leaving too little headroom for the rest of the pipeline. The archived deployment path also does not currently reproduce: a notebook shows “CUDA error: no kernel image is available for execution on the device”, and ExportOnnxRuntime.py has an unclosed parenthesis.
Why?. The core limit is semantic: the center of the visible road is not the center of the bicycle’s lane. At intersections, wide roads, turn pockets, and multi-lane sections the road mask describes several possible corridors. The 12 FPS figure has incomplete configuration records (no matching benchmark artifact for device settings, input size, or timing scope).
What I learned?. A road mask does not identify the current lane, and a convincing road overlay can hide weak lane-boundary performance.
Semantic segmentation with different BDD100K label schemes.
Why I tried it. A well-understood segmentation baseline let me test which label definitions were worth predicting.
What I implemented.
DeepLabV3-style segmentation with ResNet50 backbones and atrous spatial pyramid pooling.
BDD100K processing with configurable category selection and merging: drivable only, separate lane-border classes, and a merged lane-marking + curb class.
Training/validation metrics, image, webcam, and video inference, and SageMaker training.
What went wrong?. During testing the model, I found out this model is not well suited for real-time detection as it takes far too MS in order to compute one frame.
Why?. This model is designed for accuracy and not real-time inference.
What I learned?. The highest accuracy scores often lead to the longest inference time.
DDRNet
Checkpoint-validated on CPU
Efficient road segmentation; the main work was confirming the right weights were loaded.
DDRNet road masks. They are plausible, but the road region often covers more than the bicycle’s lane.
Why I tried it. DDRNet keeps interacting low- and high-resolution branches to balance detail and computation, which suits an embedded budget.
What I implemented.
DDRNet-23-slim and DDRNet-23 segmentation checks (DDRNet-39 definitions are present but not a completed experiment).
Separated ImageNet classification checkpoints from Cityscapes segmentation checkpoints, matched each to its architecture, and verified strict loading, handling training-only keys explicitly.
Decoding: 19-class logits, resize, argmax, standard Cityscapes road train ID.
Validation over 180 consecutive frames and nine samples spread across the video: finite outputs, checkpoint compatibility, and a decodable comparison video.
What went wrong?. Nothing failed outright. The masks are plausible, but they do not isolate the travel lane, and this was CPU validation only: no Jetson FPS and no labeled bicycle accuracy.
Why?. Cityscapes “road” is the whole road surface, not the ego lane.
What I learned?. Check checkpoint identity before judging a model: classification weights have no trained segmentation head. DDRNet’s usefulness for lane keeping is unresolved, not disproven.
Sample images and raw label masks; DDRNet test_video.ipynb
Dataset adaptation
Much of the work was understanding and converting annotations, not just downloading datasets.
CULane, TuSimple, BDD100K, and OpenLane
Implemented
Loaders, converters, and annotation viewers for datasets that describe lanes in different ways.
Why I tried it. No public dataset uses a bicycle-mounted camera. I needed to know what each dataset actually labels before training on it.
What I implemented.
Annotation viewers for TuSimple, CULane, BDD100K, OpenLane, and OpenLane-V2 (model/TuSimple.ipynb, model/CuLane.ipynb, Bdd100Test.ipynb, OpenLane*.ipynb).
CULane and TuSimple loaders for LaneATT and LaneNet/H-Net; the BDD100K LaneATT adapter; semantic lane-attribute support.
COCO and BDD100K conversions for object detection (see YOLO11 below).
Dataset
How I used it
Limitation
CULane
Main LaneATT training and evaluation; LaneNet/H-Net loaders
Vehicle-camera imagery; does not establish bicycle reliability
TuSimple
Early exploration, LaneNet infrastructure, H-Net training
Constrained highway imagery
BDD100K
Segmentation, LaneATT adaptation, lane attributes, object-detection prep
Different annotation semantics; conversion choices matter
COCO 2017
YOLO11 fine-tuning on road-relevant classes
No lane geometry or distance
OpenLane / OpenLane-V2
Annotation and multi-camera topology exploration
No completed bicycle model trained from it
Cityscapes DDRNet checkpoints
Road-segmentation experiments
A road mask is not the travel lane
Virtual KITTI metric depth
Metric-depth proof of concept
Meters unvalidated on our camera
Recorded road and bicycle video
Offline comparison, debugging, failure analysis
Mostly unlabeled
What went wrong?. Most of our own bicycle recordings have no complete manual labels, so every bicycle-camera result on this page is qualitative or a field estimate.
Why?. Public datasets come from car cameras, and each labels something different: CULane lane polylines, BDD100K marking segments and boundary objects, segmentation masks, and object boxes.
What I learned?. Dataset semantics are part of the model. The biggest missing dataset is a small, labeled set from our own camera.
Configurations, files, and additional videos
Scope notes
Local folders: CuLaneDataset/, TUSimple/, 100k_images/, 100k_json/, Bdd100Final/, Bdd100Final2/, OpenLane/
bddModels/ is an upstream reference collection of BDD100K task implementations; I did not train every model in it.
Supporting investigations
Work tied to the project’s broader goals of obstacle detection, distance estimation, and collision avoidance. My immediate focus has since narrowed to lane keeping.
YOLO11 object detection
Evaluated on COCO · runs in ROS 2
A two-class road-object detector exported to TensorRT and run as a ROS 2 node.
Why I tried it. Obstacle awareness needs object locations alongside. Easist object detection model to implement.
What I implemented.
COCO conversion to four classes (person; vehicle = car)
Nano, small, and medium models; PyTorch, ONNX, embedded-NMS export, TensorRT, and ROS 2 inference returning (1, 300, 6) boxes.
Run
Best val mAP50–95
mAP50 at same epoch
yolo11n_coco45
0.46546
0.65025
yolo11s_coco43
0.51845
0.71005
yolo11m_coco43
0.54285
0.73225
COCO validation metrics, not obstacle-avoidance performance on the bicycle.
What went wrong?. Nothing in the detector itself. This model worked perfected fine in everything.
Depth Anything V2 (monocular depth)
Proof of concept
Relative and metric monocular depth as a first step toward obstacle distance.
Relative depth on recorded footage. Colors are normalized per frame and do not show a consistent physical scale.
Why I tried it. Distance and closing speed would be needed for any stopping decision.
What I implemented.
Depth Anything V2 Small wrapper, image/video inference, ONNX export and checks (518 × 518 input), and TensorRT inference in the combined benchmark.
Metric-depth test with the Virtual-KITTI Small checkpoint, saving depth_meters_frame100.npy and metric_frame100.png.
What went wrong?. The standard model outputs relative depth, which is ordering, not meters. The metric variant ran and produced numbers, but nothing validates those distances on the bicycle camera.
Why?. A metric model trained on Virtual KITTI still needs camera and domain validation. An 80 m output range says nothing about accuracy at 80 m, and differentiating noisy distances gives very noisy closing-speed estimates.
What I learned?. Relative depth cannot support stopping decisions. The models in depthandspeedestimationresearch.md are researched options, not implemented ones.
Stereo depth with NVIDIA VPI
Partially implemented
Geometric depth from the IMX219 stereo pair as a ROS 2 node.
Why I tried it. Stereo gives depth from geometry rather than learned appearance.
What I implemented.
Approximately synchronized left/right topics, grayscale, 640 × 360, VPI disparity, fixed-point to pixel conversion, Z = fB / disparity, and published depth plus a disparity visualization.
What went wrong?. Calibration and integration are unresolved. The node uses an assumed 60 mm baseline and a focal length from approximate sensor specs rather than a calibrated CameraInfo, and it is not a full rectified stereo pipeline. Topic names do not match the rest of the system, and the depth image is smaller than the source image the detection boxes come from.
Why?. It subscribes to /left/image_raw and /right/image_raw while other nodes use /stereo/left/image_raw, and it publishes /stereo/depth_image while the object controller expects /stereo/depth. The controller does not rescale boxes before indexing the smaller depth map.
What I learned?. An implemented depth node is not usable depth until calibration, topic wiring, and coordinate scaling are verified.
Simulation and hardware
Partially implemented
Gazebo path-following tests and ESP32-S3 firmware.
Why I tried it. Simulation tests path-following logic with a known path; the firmware is what ultimately moves the steering.
What I implemented.
Older ROS 2 workspace: Xacro/URDF robot, Gazebo straight and curved road worlds, generated road/boundary/centerline meshes, manual control, and a Stanley node following a CSV centerline from odometry.
ESP32-S3 firmware on the origin/S3-code branch: AS5600 steering-angle feedback, proportional steering, throttle PWM, brake logic, manual/remote modes, I2C and serial angle commands.
What went wrong?. The simulated robot is a simplified wheeled platform, not a validated bicycle dynamics model. The firmware exists, but there is no verified bridge from current ROS 2 perception to physical steering and braking.
Why?. Path following with a known path and recovering that path from a camera are separate problems; the simulation only covers the first.
What I learned?. Controller logic can be tested independently of perception, but the connection between them is where the project currently stops.
Failure analysis
Several of the most important failures happen outside the neural network. Each case lists what should happen, what did happen, the suspected cause, and whether that cause has been verified.
1.Missing or incorrectly fitted lane boundaries
LaneNet + H-Net: the right fitted curve bends away from the visible boundary, then disappears.
Expected
Both boundaries of the bicycle’s lane are drawn whenever the markings are visible.
Observed
The right curve bends away in one frame and is missing in another. Separately, LaneATT gave usable left/right boundaries in fewer than ~25% of bicycle frames (field estimate).
Suspected cause
Unstable clustering and curve fitting for LaneNet; camera-domain shift for LaneATT.
Verified?
The instability is observed in inspected frames. Neither cause has been isolated, and there is no whole-video failure rate.
2.Valid predictions discarded during ego-lane selection
Left boundary
y = 400–700
Right boundary
y = 400–700
Unrelated prediction
y = 50–150
Selector intersects every prediction's vertical range
Empty shared interval → valid pair discarded → no lane output
Reproducible selector test; the third prediction is outside the valid pair's vertical range.
Expected
The selector returns the left/right pair (and a centerline) whenever the model predicts it, and still returns something useful with one boundary.
Observed
A valid pair returns two boundaries and a centerline. Adding an unrelated third lane with no vertical overlap makes the same function return no lanes. A single boundary also returns no lanes.
Suspected cause
get_ego_lanes2 first computes a vertical interval common to all predicted lanes, and only then chooses the closest left/right candidates. One unrelated lane can empty that interval.
Verified?
Yes, reproduced in a small test. How often this happens on real bicycle footage has not been measured. Confidence hysteresis and lane-width-based boundary synthesis exist elsewhere in the repository but are not active in this selector.
Deployment went from PyTorch to ONNX to ONNX Runtime to TensorRT FP16 engines, and then to a Torch-free runtime on the Jetson. Along the way I added per-stage timing, render and no-render comparisons, sequential and separate-process runs, power-mode comparisons, and numerical parity checks. TensorRT FP16 made LaneATT about 3.7× faster than ONNX Runtime FP32 under the same benchmark scope.
Model
Runtime
Precision
Input
Power mode
Throughput
Timing includes
LaneATT
ONNX Runtime
FP32
640 × 360
MAXN_SUPER
7.33 FPS
Video read + preprocess + engine
LaneATT
TensorRT
FP16
640 × 360
MAXN_SUPER
27.32 FPS
Same scope as above
YOLO11n
ONNX Runtime
FP32
640 × 640
MAXN_SUPER
26.98 FPS
Single-model benchmark
YOLO11n
TensorRT
FP16
640 × 640
MAXN_SUPER
39.94 FPS
Single-model benchmark
LaneATT + YOLO + depth
TensorRT, sequential
See record
Per model
15 W
9.88 FPS
Separate three-model experiment
LaneATT + YOLO + depth
TensorRT, sequential
See record
Per model
MAXN_SUPER
18.95 FPS
Same experiment
Recorded Jetson Orin Nano Super benchmarks. Input sizes are each model’s standard exported input. None of these is a camera-to-actuator ROS 2 measurement.
Engine time is not pipeline time
Stage
Time per frame
Reading the video
≈ 40.9 ms
Two TensorRT engines
≈ 42.8 ms
Rendering
≈ 24.3 ms
Benchmark_6.json: where one frame’s time went in a measured two-engine pipeline (8.72 FPS overall). An engine-only headline would hide the rest.
Field measurements and estimates
HybridNets at ≈ 12 FPS after TensorRT is my field measurement. No matching benchmark record exists for its device configuration, input size, or timing scope.
LaneATT producing usable boundaries in fewer than ≈ 25% of bicycle frames is a field estimate, pending a labeled evaluation.
Separate concurrent processes reached ≈ 35.75 LaneATT FPS and ≈ 43.07 YOLO FPS. These cannot be added: independent processes may be working on different frames, so this is not synchronized perception at the combined rate.
Some sweeps in the benchmark documentation are incomplete. Pending results remain pending.
Deployment notes
Torch-free inference: a CUDA-runtime wrapper around TensorRT 10.3 manages GPU buffers directly, so PyTorch is not needed on the Jetson.
Parity checks compare TensorRT and ONNX outputs numerically.
LaneATT exports raw proposals; confidence filtering, lane NMS, decoding, and ego-lane selection run afterward. YOLO exports with NMS embedded. A running engine is not a correct final output.
Engines are built for their target: an engine built on an A100 is not a Jetson artifact.
The Orin Nano has no DLA, so designs that split models between DLA and GPU do not apply. Advertised TOPS is not measured application throughput.
Connecting perception to steering
Running the models does not by itself keep the bicycle in its lane. Lane keeping also needs lane selection, path geometry, timing, and the connection to hardware, and each of those is at a different stage.
Camera input
IMX219 stereo driver (1280 × 720, 60 FPS requested), USB cameras, or recorded-video replay
Runs on the Jetson
→
Model predictions
LaneATT TensorRT raw proposals; YOLO11 boxes
Runs on the Jetson
→
Lane selection
Confidence 0.2, lane NMS 50, top 4, then get_ego_lanes2
Runs on the Jetson
→
Path geometry
Center path, heading, and cross-track error (angle.py, road_guidance.py)
Tested offline
→
Steering
Stanley-style estimate → ESP32-S3 steering and throttle
Needs integration and validation
Green: runs on the Jetson in ROS 2. Blue: implemented and tested offline on recorded video. Dashed amber: code exists (ESP32-S3 firmware by Rayan) but the perception-to-actuator connection has not been verified.
ROS 2 nodes
Package
Purpose
Current state
imx219_83_jetson_driver
Stereo capture
C++/GStreamer left/right images and CameraInfo
frame_publisher
Recorded-video replay
Publishes to /stereo/left/image_raw; loops at EOF
laneatt2
TensorRT lane perception
Publishes normalized left, right, and middle lane points; uses the Sept ResNet34 engine
Median depth in the box; < 8 ft slow, 8–10 hold, > 10 speed up
center_lane
Center-lane controller
Incomplete stub
video_test
Camera/debug display
Image conversion and callback timing
Why model inference has not yet become lane keeping?
Timestamps and freshness. Lane and detection messages are consumed as “latest received”. Detection messages carry no source-image timestamp, and nothing rejects stale outputs.
Coordinate conversion. LaneATT publishes normalized points from a 640 × 360 input; YOLO letterboxes to 640 × 640; stereo depth is smaller than the source image. Each consumer must convert consistently, and the object controller does not rescale boxes before indexing depth.
Missing boundaries. With one visible boundary the active selector returns nothing. Lane-width-based synthesis exists in other code but is not active here.
Preprocessing mismatch. Offline video inference can pad frames onto a larger canvas; the ROS 2 node resizes directly. Results from one path do not transfer automatically to the other.
Steering assumptions. angle.py assumes a 90° horizontal field of view, a 3.7 m lane width, and a fallback speed. These are assumptions, not calibration or measured telemetry.
Undefined fallback. object_tracker_controller defaults to “speed up” when no valid distance is available. Behavior when lanes or distances disappear still has to be defined.
Launch files. full_system.launch.py currently launches USB camera nodes only, despite its name. Requested camera rates (60 or 30 FPS) are not measured delivery rates.
What I learned
Strong dataset results do not establish bicycle-camera reliability.
LaneATT reached 0.785 CULane validation F1 but gave usable boundaries in fewer than ~25% of bicycle frames (field estimate). See evidence
A road mask does not necessarily identify the current lane.
HybridNets road IoU ≈ 0.84 vs. lane IoU ≈ 0.26; DDRNet masks span adjacent lanes. See evidence
Postprocessing can lose useful model predictions.
An unrelated third lane makes get_ego_lanes2 reject a valid left/right pair. See evidence
Fast inference does not establish fresh, synchronized output.
The same overlay persisted from frame 4000 to 6800; video reading and rendering took more time than the engines in Benchmark_6. See evidence
Relative depth does not provide validated stopping distances.
Depth Anything V2 relative output is ordering only; the metric variant’s meters are unvalidated on our camera. See evidence
Preprocessing and configuration are part of the model.
BGR order, padding vs. resize, checkpoint identity (DDRNet), and a scheduler that decays once per epoch instead of once per run (BDD100K). See evidence
Proposed next experiment
Proposed, not yet done
Replay a fixed set of labeled bicycle frames through the exact deployed pipeline and measure where useful lane information is lost.
Label the ego-lane boundaries on a small, fixed set of bicycle-camera frames covering straight roads, curves, shadows, one-sided markings, and intersections.
Run them through the same preprocessing and engine as the ROS 2 node.
Record the result at each stage: raw proposals, after confidence filtering and NMS, after ego-lane selection, the published ROS 2 message (with source timestamp), and the drawn overlay.
Report, per stage, how often a correct boundary survives, so model errors, selection errors, and delivery errors are measured separately.
Where I would value guidance
Should this diagnosis come before any further model changes or retraining?
What evidence would justify the next controlled lane-keeping test on the bicycle, for example a minimum per-stage boundary rate, an end-to-end latency bound, and stale-output rejection?
Is a simpler geometry approach, such as tracking one reliable boundary with a lane-width offset, a reasonable target at bicycle speeds?
Research materials and credits
Poster and code
Research poster PDF: original file being located
Code repository: link to be added
Experiment logs and records
LaneATT training logs (Aug2 and Sept 10 runs); BDD100K run log
HybridNets metrics.jsonl, meter2.jsonl
LaneNet inference_runs.json
Jetson benchmark JSON records (e.g. Benchmark_6.json) and benchmark documentation
DDRNet MODEL_VALIDATION.md and report.json
YOLO11 training results (yolo11n_coco45, yolo11s_coco43, yolo11m_coco43)
openpilot (upstream checkout; not ported to the bicycle)
OpenLane-V2 devkit; LaneTCA upstream code
RONELDv2 lane tracking
Intelligent Driver Model survey, a car-following control reference
Metric-depth and speed-estimation survey
bddModels/ upstream BDD100K implementations; an earlier YOLOPv2 reference
Team
Aman Nindra · Computer vision, dataset preparation, model training, inference optimization, and perception integration.
Rayan · Bicycle robotics and control hardware, including the ESP32-S3 controller, steering feedback, and 250 W motor system.
LaneATT, LaneNet, H-Net, HybridNets, DeepLab, DDRNet, LaneTCA, YOLO11, and Depth Anything V2 are published research systems adapted here for training, evaluation, and deployment experiments; their architectures are not claimed as original work.