Skip to content
View syedhamzamohiuddin's full-sized avatar

Block or report syedhamzamohiuddin

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
syedhamzamohiuddin/README.md

Hi, I'm Syed Hamza Mohiuddin πŸ‘‹

Computer Vision Engineer | Systems & Edge AI

I am an algorithmic and systems-focused Computer Vision Engineer with a holistic grasp of the vision stack. My expertise bridges the gap between classical computer vision techniques (image processing, geometric reasoning, and tracking), deep learning architectures, and high-performance production pipelines: from custom training loops and multi-process system design to model optimization and edge deployment (TensorRT, Qualcomm AI Hub, ONNX, TFLite).


πŸ† Key Highlights

  • IEEE LPCVC 2026 (CVPR Workshop) – Track 2: Ranked 8th/38 globally in Video Action Recognition. Optimized R2+1D for Qualcomm Dragonwing IQ-9075 under a strict <34ms latency budget, overcoming severe class imbalance and label noise through rigorous data auditing.
  • IEEE LPCVC 2026 (CVPR Workshop) – Track 1: Ranked 12th/56 globally in Open-World Retrieval. Optimized MobileCLIP/ViT-B/16 for Qualcomm XR2 Gen 2, engineering a custom model wrapper to resolve tokenizer discrepancies (Causal vs. Bidirectional) for on-device inference.
  • High-Performance Systems Architecture: Redesigned a monolithic, single-process 30 FPS dual-camera pipeline into a 4-process parallel architecture using Zero-Copy IPC (multiprocessing.shared_memory) and OpenCV-CUDA. Achieved over 3x throughput (30 β†’ 100+ FPS per camera stream) by resolving GUI latency and CPU-bound preprocessing bottlenecks.

πŸ› οΈ Core Expertise

  • Systems Architecture & Optimization: High-Performance Multiprocessing, Zero-Copy IPC (Shared Memory), GPU Acceleration (OpenCV-CUDA), TensorRT, Qualcomm AI Hub (QNN), ONNX, INT8 Quantization, Structured Pruning.
  • Deep Learning & DL-based CV: Object Detection (YOLO v5-v11), Segmentation (U-Net), Tracking (ByteTrack, TransReID), Pose Estimation, Vision Transformers (ViT/VLMs), ReID, Video Action Recognition (Temporal Modeling), Multi-modal Retrieval (CLIP, MobileCLIP).
  • Classical & Algorithmic CV: Image Processing (Morphology, CLAHE, Canny, Contours), Geometric Reasoning (Zhang's Calibration, Homography, ArUco, scipy.optimize), Tracking (Kalman Filtering, Optical Flow, Background Subtraction).
  • Languages & Tools: Python (Expert), PyTorch, TensorFlow, C++ (CUDA Certified), Netron Model Surgery, Reverse Engineering, Model Benchmarking (clip_benchmark).

πŸ’Ό Experience

  • Swift Vision Pvt. Ltd. | Computer Vision Intern β†’ Computer Vision Engineer β†’ Computer Vision Research Engineer | Apr 2025 - Jan 2026 (Promoted to R&D ownership within 4 months based on execution and technical leadership).
  • Folio3 Pvt. Ltd. | AI Intern | July 2023 - Sep 2023

πŸš€ Featured Projects

High-Performance Multi-Camera Vision Pipeline (30 FPS β†’ 100+ FPS)

  • Role: Lead Systems Architect (Swift Vision)
  • Challenge: Client's monolithic 30 FPS dual-camera sports analytics system was bottlenecked by sequential CPU-bound processing and OpenCV GUI latency. The client initially misdiagnosed the issue as a model inference problem.
  • Solution:
    • Diagnosed the architectural flaw and argued against the TensorRT misdiagnosis.
    • Redesigned the system into a 4-process parallel architecture: two camera logic processes (GPU preprocessing + YOLO inference), an update/state loop, and a main GUI process.
    • Implemented Zero-Copy IPC using multiprocessing.shared_memory to eliminate data copying overhead.
    • Offloaded all image preprocessing (resize, warp, blur, dilation) to OpenCV-CUDA.
  • Result: Achieved over 3x throughput (30 β†’ 100+ FPS per stream, with peaks near 120) and 120 FPS GUI rendering. Received direct praise from the client for the architectural solution.

IEEE LPCVC 2026 (CVPR Workshop) – Track 2: Video Action Recognition

  • Result: Ranked 8th/38 globally.
  • Challenge: 92-class fitness video dataset with severe class imbalance (some classes <150 examples) and significant label noise.
  • Solution & Approach: Modified the training pipeline to log per-class train/val accuracy and individual misclassifications, manually auditing 4,500+ problematic instances. Designed 7 separate augmentation and cropping strategies based on class-specific confusion analysis. Diagnosed a Qualcomm hardware compiler tiling failure and pivoted to a 16-frame architecture to meet a <34ms latency budget.
  • Outcome: Improved accuracy from 91.44% to 93.24% through rigorous data curation and training rule refinement.

πŸ”§ Open Source: Ultralytics YOLO

  • TensorRT validation fix (PR #21592, merged) β€” Root-caused a yolo val failure on TensorRT .engine models (the validator read a batch_size attribute that AutoBackend never set) and worked through maintainer review; the merged fix simplified the validator's batch-size logic.
  • Structured Pruning Engine (PR #21977, docs PR #22438) β€” Native PyTorch structured pruning for YOLOv8 detection models.
    • Global ratio or per-layer YAML configuration; norm-based channel importance; mask propagation through Conv, BatchNorm, Bottleneck, C2f, SPPF, Concat and both Detect-head towers; group-aware handling of grouped/depthwise convolutions (the reason torch.nn.utils.prune wasn't enough).
    • ~1,000 lines of pruning code plus a 1,400-line test suite (43 test functions) covering the prune β†’ save β†’ load β†’ retrain β†’ ONNX export round trip.
    • Functional demo (YOLOv8s, COCO128, ONNX): 22.6 β†’ 11.6 MB (~49% smaller), 17.1 β†’ 12.0 ms (~30% faster). A demonstration that the pipeline works end to end, not an accuracy benchmark; retraining is needed to recover accuracy.
    • Status: scoped with maintainer guidance, CI green and branch up to date; both PRs were closed by the stale bot before code review. Code is on my fork: feature/prune-functionality and docs/pruning.

Paper Implementations

Reimplemented from the original papers, first principles.

  • Transformer (Attention Is All You Need) β€” Complete preprocessing and training pipeline, shared source-target vocabulary, byte-pair encoding, three-way weight tying, padding and causal masking, full encoder-decoder architecture.
  • U-Net (Biomedical Segmentation) β€” Pixel-wise weight maps and elastic deformation, optimized for TPU-accelerated training using a TFRecord-based data pipeline.

Other Projects

  • Real-Time Liquid Quality Inspection β€” Automated inspection pipeline for transparent bottles on a conveyor using backlit imaging; morphological segmentation, CLAHE, Canny edge detection, contour analysis, and Kalman filter tracking.
  • Clinical Caries Detection β€” Customized U-Net for dental radiography segmentation, high-precision boundary detection.
  • Traffic Analytics Pipeline β€” Automated traffic analysis using optical flow, thresholding, and multi-object tracking (MOT) to monitor vehicle flow and density.
  • Deep Probabilistic Generative Models β€” VAEs for facial generation and RealNVP (Normalizing Flows) for high-dimensional image synthesis on LSUN.
  • Bayesian CNN (Uncertainty Quantification) β€” Captures aleatoric and epistemic uncertainty in digit classification, beyond point-estimate predictions.
  • Bayesian Earthquake Forecasting β€” Bayesian AR(3) model in R for global seismic activity: model order selection, posterior inference, prior sensitivity, mixture AR model comparison..

πŸŽ₯ Video Demonstrations

A collection of visual demos from early project work showcasing classical CV, tracking, and basic detection pipelines.


πŸ“– Technical Foundations & Learning Roadmap

  • Foundational Exposure: Studied early chapters of Hartley & Zisserman's Multiple View Geometry to understand the mathematical principles behind production geometry tasks.
  • Applied Practice: Translated these geometric foundations into working code, including Zhang's planar calibration and marker-aware ArUco rectification using scipy.optimize.
  • Current Learning Focus: Currently building depth in Kalman Filtering and C++ for high-performance geometric vision.
  • Upcoming Study Plan: Next steps include 3D vision libraries (PCL) and advanced coursework (CS231A) to deepen 3D perception expertise.

πŸŽ“ Education

  • BS Computer Science | IBA Karachi (Best Paper Nomination, ICETST 2022)

Certifications & Coursework

Click to expand β€” 13 certifications across ML theory, CV, and deep learning

Core Certifications

Coursera Specializations

Other


πŸ“« Connect with Me

Pinned Loading

  1. ultralytics ultralytics Public

    Forked from ultralytics/ultralytics

    Ultralytics YOLO πŸš€

    Python

  2. unet-paper-reimplementation unet-paper-reimplementation Public

    Faithful implementation of the original U-Net architecture with weight maps, elastic deformation, and full training pipeline β€” based on the 2015 paper.

  3. neat-flappy neat-flappy Public

    Java

  4. bayesian-earthquake-forecast bayesian-earthquake-forecast Public

    R

  5. Papers Papers Public

    Paper Implementations

    Jupyter Notebook 1