I am an algorithmic and systems-focused Computer Vision Engineer with a holistic grasp of the vision stack. My expertise bridges the gap between classical computer vision techniques (image processing, geometric reasoning, and tracking), deep learning architectures, and high-performance production pipelines: from custom training loops and multi-process system design to model optimization and edge deployment (TensorRT, Qualcomm AI Hub, ONNX, TFLite).
- IEEE LPCVC 2026 (CVPR Workshop) β Track 2: Ranked 8th/38 globally in Video Action Recognition. Optimized R2+1D for Qualcomm Dragonwing IQ-9075 under a strict <34ms latency budget, overcoming severe class imbalance and label noise through rigorous data auditing.
- IEEE LPCVC 2026 (CVPR Workshop) β Track 1: Ranked 12th/56 globally in Open-World Retrieval. Optimized MobileCLIP/ViT-B/16 for Qualcomm XR2 Gen 2, engineering a custom model wrapper to resolve tokenizer discrepancies (Causal vs. Bidirectional) for on-device inference.
- High-Performance Systems Architecture: Redesigned a monolithic, single-process 30 FPS dual-camera pipeline into a 4-process parallel architecture using Zero-Copy IPC (
multiprocessing.shared_memory) and OpenCV-CUDA. Achieved over 3x throughput (30 β 100+ FPS per camera stream) by resolving GUI latency and CPU-bound preprocessing bottlenecks.
- Systems Architecture & Optimization: High-Performance Multiprocessing, Zero-Copy IPC (Shared Memory), GPU Acceleration (OpenCV-CUDA), TensorRT, Qualcomm AI Hub (QNN), ONNX, INT8 Quantization, Structured Pruning.
- Deep Learning & DL-based CV: Object Detection (YOLO v5-v11), Segmentation (U-Net), Tracking (ByteTrack, TransReID), Pose Estimation, Vision Transformers (ViT/VLMs), ReID, Video Action Recognition (Temporal Modeling), Multi-modal Retrieval (CLIP, MobileCLIP).
- Classical & Algorithmic CV: Image Processing (Morphology, CLAHE, Canny, Contours), Geometric Reasoning (Zhang's Calibration, Homography, ArUco,
scipy.optimize), Tracking (Kalman Filtering, Optical Flow, Background Subtraction). - Languages & Tools: Python (Expert), PyTorch, TensorFlow, C++ (CUDA Certified), Netron Model Surgery, Reverse Engineering, Model Benchmarking (
clip_benchmark).
- Swift Vision Pvt. Ltd. | Computer Vision Intern β Computer Vision Engineer β Computer Vision Research Engineer | Apr 2025 - Jan 2026 (Promoted to R&D ownership within 4 months based on execution and technical leadership).
- Folio3 Pvt. Ltd. | AI Intern | July 2023 - Sep 2023
- Role: Lead Systems Architect (Swift Vision)
- Challenge: Client's monolithic 30 FPS dual-camera sports analytics system was bottlenecked by sequential CPU-bound processing and OpenCV GUI latency. The client initially misdiagnosed the issue as a model inference problem.
- Solution:
- Diagnosed the architectural flaw and argued against the TensorRT misdiagnosis.
- Redesigned the system into a 4-process parallel architecture: two camera logic processes (GPU preprocessing + YOLO inference), an update/state loop, and a main GUI process.
- Implemented Zero-Copy IPC using
multiprocessing.shared_memoryto eliminate data copying overhead. - Offloaded all image preprocessing (resize, warp, blur, dilation) to OpenCV-CUDA.
- Result: Achieved over 3x throughput (30 β 100+ FPS per stream, with peaks near 120) and 120 FPS GUI rendering. Received direct praise from the client for the architectural solution.
- Result: Ranked 8th/38 globally.
- Challenge: 92-class fitness video dataset with severe class imbalance (some classes <150 examples) and significant label noise.
- Solution & Approach: Modified the training pipeline to log per-class train/val accuracy and individual misclassifications, manually auditing 4,500+ problematic instances. Designed 7 separate augmentation and cropping strategies based on class-specific confusion analysis. Diagnosed a Qualcomm hardware compiler tiling failure and pivoted to a 16-frame architecture to meet a <34ms latency budget.
- Outcome: Improved accuracy from 91.44% to 93.24% through rigorous data curation and training rule refinement.
- TensorRT validation fix (PR #21592, merged) β Root-caused a
yolo valfailure on TensorRT.enginemodels (the validator read abatch_sizeattribute thatAutoBackendnever set) and worked through maintainer review; the merged fix simplified the validator's batch-size logic. - Structured Pruning Engine (PR #21977, docs PR #22438) β Native PyTorch structured pruning for YOLOv8 detection models.
- Global ratio or per-layer YAML configuration; norm-based channel importance; mask propagation through Conv, BatchNorm, Bottleneck, C2f, SPPF, Concat and both Detect-head towers; group-aware handling of grouped/depthwise convolutions (the reason
torch.nn.utils.prunewasn't enough). - ~1,000 lines of pruning code plus a 1,400-line test suite (43 test functions) covering the prune β save β load β retrain β ONNX export round trip.
- Functional demo (YOLOv8s, COCO128, ONNX): 22.6 β 11.6 MB (~49% smaller), 17.1 β 12.0 ms (~30% faster). A demonstration that the pipeline works end to end, not an accuracy benchmark; retraining is needed to recover accuracy.
- Status: scoped with maintainer guidance, CI green and branch up to date; both PRs were closed by the stale bot before code review. Code is on my fork:
feature/prune-functionalityanddocs/pruning.
- Global ratio or per-layer YAML configuration; norm-based channel importance; mask propagation through Conv, BatchNorm, Bottleneck, C2f, SPPF, Concat and both Detect-head towers; group-aware handling of grouped/depthwise convolutions (the reason
Reimplemented from the original papers, first principles.
- Transformer (Attention Is All You Need) β Complete preprocessing and training pipeline, shared source-target vocabulary, byte-pair encoding, three-way weight tying, padding and causal masking, full encoder-decoder architecture.
- U-Net (Biomedical Segmentation) β Pixel-wise weight maps and elastic deformation, optimized for TPU-accelerated training using a TFRecord-based data pipeline.
- Real-Time Liquid Quality Inspection β Automated inspection pipeline for transparent bottles on a conveyor using backlit imaging; morphological segmentation, CLAHE, Canny edge detection, contour analysis, and Kalman filter tracking.
- Clinical Caries Detection β Customized U-Net for dental radiography segmentation, high-precision boundary detection.
- Traffic Analytics Pipeline β Automated traffic analysis using optical flow, thresholding, and multi-object tracking (MOT) to monitor vehicle flow and density.
- Deep Probabilistic Generative Models β VAEs for facial generation and RealNVP (Normalizing Flows) for high-dimensional image synthesis on LSUN.
- Bayesian CNN (Uncertainty Quantification) β Captures aleatoric and epistemic uncertainty in digit classification, beyond point-estimate predictions.
- Bayesian Earthquake Forecasting β Bayesian AR(3) model in R for global seismic activity: model order selection, posterior inference, prior sensitivity, mixture AR model comparison..
A collection of visual demos from early project work showcasing classical CV, tracking, and basic detection pipelines.
- View the full project demo playlist on YouTube (Includes: Water Quality Inspection, Vehicle Detection, Pose Estimation, Pedestrian Tracking, Multi-Car Tracking, Liquid Fill Estimation, and more).
- Foundational Exposure: Studied early chapters of Hartley & Zisserman's Multiple View Geometry to understand the mathematical principles behind production geometry tasks.
- Applied Practice: Translated these geometric foundations into working code, including Zhang's planar calibration and marker-aware ArUco rectification using
scipy.optimize. - Current Learning Focus: Currently building depth in Kalman Filtering and C++ for high-performance geometric vision.
- Upcoming Study Plan: Next steps include 3D vision libraries (PCL) and advanced coursework (CS231A) to deepen 3D perception expertise.
- BS Computer Science | IBA Karachi (Best Paper Nomination, ICETST 2022)
Click to expand β 13 certifications across ML theory, CV, and deep learning
Core Certifications
- Google TensorFlow Developer Certificate
- NVIDIA β Accelerated Computing (CUDA C/C++)
- NVIDIA β Building Real-Time Video AI Applications
Coursera Specializations
- Machine Learning β Stanford University
- Deep Learning β DeepLearning.AI
- First Principles of Computer Vision β Columbia University
- Computer Vision for Engineering & Science β MathWorks
- Image Processing for Engineering & Science β MathWorks
- TensorFlow 2 for Deep Learning β Imperial College London
- Mathematics for Machine Learning β Imperial College London
- Statistics with Python β University of Michigan
- Bayesian Statistics β UC Santa Cruz
Other
- Email: syedhamza097@gmail.com
- LinkedIn: linkedin.com/in/syedhamzamohiuddin
- Kaggle: kaggle.com/hamzamohiuddin
- Medium: medium.com/@syedhamza097