You're seeing this page as if you were . The main menu is still yours, though. Exit from immersion
Mohammed GuermalMG

Mohammed Guermal

Research Engineer ML/Computer Vision

€500/day
Lyon, FR
3-7 years

Average response time: 1 hour

About Mohammed

I am a Computer Vision and Artificial Intelligence expert with a PhD in Computer Vision from Inria (France), followed by a postdoctoral position at the University of Luxembourg. I have several years of experience developing state-of-the-art AI solutions in both academic research and industry.

My expertise covers Computer Vision, Deep Learning, Multimodal AI, Vision-Language Models (VLMs), and Generative AI. I have designed and deployed AI systems for object detection, image segmentation, tracking, video understanding, anomaly detection, satellite imagery, and real-time video analytics.

I work with modern deep learning frameworks including PyTorch, Hugging Face Transformers, ONNX, TensorRT, and OpenCV, and have extensive experience building end-to-end machine learning pipelines—from data preparation and model training to optimization and deployment on edge and cloud platforms.

Whether you need a research prototype, a production-ready AI solution, or expert consulting on cutting-edge AI technologies, I can help transform complex ideas into reliable, high-performance systems.
  • Arabic

    Native or bilingual

  • English

    Fluent

  • French

    Native or bilingual

  • Spanish

    Basic

Can work on-site
Lyon (up to 50km), Paris (up to 50km), Nice (up to 50km)

Experience

  • Poseidon
    Research Engineer in Computer Vision
    SPORTS
    January 2026 - Today (7 months)
    Paris, France
    Building a computer vision system for real-time drowning detection in pools, from swimmer tracking to alert generation Explored and benchmarked detection/tracking pipelines for swimmers: YOLO, SAM3, SORT, DeepSORT, Ultralytics -- evaluating trade offs between tracking robustness and real-time constraints in a challenging visual environment (water, occlusion, reflections) Designed video classification models for real-time drowning-alert detection using 3D CNNs and Transformer architectures Led the creation of a purpose-built swimmer-tracking dataset, since no public dataset covers pool-based drowning scenarios Researching how vision-language models (Qwen, LAVAD, LLaVA-NeXT) can be adapted as a downstream reasoning layer on top of raw detections, to add contextual interpretation to drowning alerts Applying temporal action-anticipation techniques (LSTR, TesTra) -- extending the online action-anticipation research from my PhD (see JOADAA, WACV 2023) -- to predict drowning risk before it fully develops, rather than only detecting it after the fact
    Pytorch Python Streamlit Git Docker
  • Veesion
    Research Engineer ML/Computer Vision
    January 2025 - January 2026 (1 year)
    Paris, France
    Built neural network models for theft detection in retail video streams, combining 3D CNNs and Transformer-based architectures for anomaly/action recognition -- an industry application closely related to my anomaly-forecasting research (see "Guess Future Anomalies from Normalcy", WACV 2025) Designed multimodal models fusing skeleton, RGB, and text modalities, extending my PhD work on multimodal fusion (see CM3T, WACV 2025) to capture complex movements and video context beyond a single modality Explored self-supervised pretraining to reduce dependence on annotated data: experimented with SimCLR, MAE, VideoMAE, and V-JEPA to learn general-purpose video representations before fine-tuning on theft-detection tasks Investigated contrastive learning for video-text alignment -- SigLIP, CLIP, X-CLIP -- to identify suspicious behaviors and thefts across long, weakly-labeled video sequences Applied knowledge distillation to compress models for embedded/edge deployment on in-store storage servers, reducing memory footprint and latency (up to 60% in production tests) without materially sacrificing detection accuracy Integrated large vision-language models (Qwen via Hugging Face, LLaVA-NeXT Video) as a contextual-reasoning layer on top of detection outputs, to interpret ambiguous actions rather than relying on classification alone Ran full training pipelines on AWS GPU infrastructure, from data/resource management to deploying prototypes for real-world validation
    Amazon Web Services Python Traitement d'image LLM vllm
  • University of Luxembourg
    Research Assistant
    January 2024 - January 2025 (1 year)
    Luxembourg
    Developed and optimized human pose estimation algorithms to detect and track human movements in video, targeting applications in spatial/aerospace environments Designed unsupervised domain adaptation methods to transfer object detection models to challenging, unfamiliar visual environments without annotated data -- addressing the common gap between lab-trained models and deployment conditions Investigated video-based pose estimation across varying lighting and visibility conditions to improve robustness beyond standard benchmark settings Proposed new algorithmic approaches for pose estimation under limited-data settings, where annotation is expensive or unavailable Collaborated across disciplines with AI researchers and aerospace industry experts to integrate research prototypes into real-world applications

Recommendations

Be the first to recommend Mohammed

Help this freelancer shine by sharing your experience working together.

These freelancer profiles also match your criteria

AgathaA

Agatha Frydrych

Backend Java Software Engineer

4.7

(3)

2

BaptisteB

Baptiste Duhen

Fullstack developer

4.6

(4)

5

AmedA

Amed Hamou

Senior Lead Developer

4

(2)

7

AudreyA

Audrey Champion

Web developer

4.3

(3)

4

Education

  • PhD in Computer Science
    AI Université Côte d'Azur
    PhD in Computer Science
  • M.Sc. in Electronics
    Enseirb-Matmeca
    2020
    M.Sc. in Electronics

Skill set

Categories