Computer Vision: market
Computer Vision Engineer (CV) — sister discipline to NLP in the ML family, focused on image and video processing (not text). A mature field (since the 1960s), re-assembled by deep learning in 2012 (AlexNet) → transformers 2020–2021 (ViT) → foundation models 2023–2024 (DINOv2 / SAM / CLIP) → the generative boom 2022–2026 (Stable Diffusion / FLUX). Task focus: image classification, object detection (YOLO + Detectron2 + DETR), semantic / instance / panoptic segmentation (SAM + Mask R-CNN), OCR (PaddleOCR + EasyOCR + Tesseract legacy), face recognition + biometrics, video understanding (action recognition + tracking — ByteTrack / DeepSORT), multimodal (CLIP + GPT-4V + Claude vision + Gemini Vision), 3D vision (NeRF + Gaussian Splatting + structure-from-motion), generative image / video (Stable Diffusion / FLUX / Sora family), edge deployment (TensorRT + ONNX + CoreML + OpenVINO). Role family: Computer Vision Engineer (general — production CV pipelines), Senior CV Engineer (multi-task CV ownership + custom architectures), ML Research CV (academic-track — papers at CVPR / ICCV / ECCV — overlap with Research Engineer / Scientist), 3D Vision Engineer (NeRF / Gaussian Splatting / point cloud — rising 2024+), Generative AI Engineer (Image / Video) (Stable Diffusion / FLUX / Runway / Pika / Kling specialisation — overlap with ai-engineer), Edge CV Engineer (TensorRT + mobile deployment specialisation), Robotics Vision Engineer (SLAM + perception for robots / autonomous systems). Stack 2026: Python (mostly — some edge components in C++/Rust). PyTorch + torchvision (foundation — 90%+ of CV research on PyTorch in 2026), OpenCV (classical CV — still huge for preprocessing + traditional algorithms), Pillow (image manipulation). Self-supervised + foundation models: DINOv2 (Meta — best vision backbone 2024), CLIP + OpenCLIP + SigLIP (vision-language). Classification backbones: ViT family, Swin Transformer, ConvNeXt (modern CNN), EfficientNet, timm (PyTorch Image Models — pretrained models, must-library). Data augmentation: albumentations (industry standard — fast + comprehensive), torchvision.transforms.v2 (modern), RandAugment + AugMix. Multimodal LLM: GPT-4V, Claude 3.5 Sonnet vision, Gemini 1.5 Pro / 2.0, LLaVA + Qwen 2 VL + InternVL (open-source vision-language). OCR: PaddleOCR (Baidu — multilingual leader 2026), EasyOCR, Tesseract (legacy still big), Docling (IBM — modern doc understanding 2024+), Surya. Face recognition: InsightFace (open-source SOTA 2026 — ArcFace + RetinaFace), DeepFace, face_recognition. Russian face-rec: NtechLab FaceNGN / VisionLabs Luna / Tevian. Video tracking: ByteTrack (fastest 2026), DeepSORT (classic), BoT-SORT, StrongSORT. Video models: VideoMAE, X-CLIP, InternVideo. 3D vision: NeRF (original 2020), Instant-NGP (NVIDIA — 100× faster), Gaussian Splatting (3DGS) (rising 2023–2026 — photorealistic rendering), Nerfstudio (unified framework), Open3D (point cloud processing), PyTorch3D (Meta). Generative video: Runway Gen-3, Pika 1.5, Kling AI (Kuaishou), Sora (OpenAI), AnimateDiff + SVD (open-source). Russian generative: Sber Kandinsky 3. UIs: ComfyUI (node-based — research favourite), Automatic1111 (web UI legacy). Fine-tuning generative: LoRA + DreamBooth for Stable Diffusion / FLUX. Inference serving: Triton Inference Server (NVIDIA — multi-framework + dynamic batching), BentoML, TorchServe. Annotation: CVAT (OpenCV — industry standard for detection / segmentation), Label Studio, V7 Darwin (commercial), Roboflow (popular for quick prototyping). Datasets: COCO + ImageNet + OpenImages + LAION-5B + Objects365 + Visual Genome. According to Zorky CRM, 29 active openings with explicit CV specifics (the real pool is much wider — many CV roles are classified as general ML Engineer / Backend / Robotics). Median not published. Top stack: python, pytorch, c++, tensorflow, opencv. 24% remote.
The Computer Vision market currently has 29 open roles, 7 of them freshly observed. Median salary not published. Observed candidate pool — not published.
24% of CV Engineer jobs are remote or hybrid. CV work fully cloud-based standard. Outsourcing shops — almost always remote. Russian banks — hybrid/office. Autonomous vehicle companies — research roles often remote, deployment requires on-site (Pittsburgh / Mountain View / Palo Alto). Generative AI companies — full-remote standard. Big Tech CV — hybrid-standard.
⚠ salary known for 0 of 29 jobs; remote share from 26 with a stated format; 7 counted as fresh observations; trend and hiring difficulty are not shown
Demand and observed supply
| Open demand | 29 |
| Observed supply | — not published |
⚠ candidates matching a vacancy are not counted yet: demand and the observed pool are shown
Demand geography
| country | jobs |
|---|---|
| KZ | 5 |
| US | 5 |
| UA | 2 |
| CH | 2 |
| LT | 1 |
| MA | 1 |
| FR | 1 |
| PL | 1 |
| DE | 1 |
| JP | 1 |
The leader by CV Engineer job count is Russia (0 positions). Poland — CV-friendly EU hub. Germany — Berlin AI + Munich automotive (BMW / Mercedes AI / Bosch / Continental autonomous teams). UK — London (Wayve autonomous + Tractable). USA — Bay Area + Pittsburgh (autonomous vehicle clusters) + Boston (Robotics MIT region). Huge international remote via autonomous vehicle companies (Tesla / Waymo / Cruise / Wayve / Pony.ai / Zoox / Aurora) + generative AI (Stability / Black Forest Labs / Runway / Pika / Kling) + Big Tech CV + Y Combinator CV startups.
⚠ job counts only: salary by country is not published
Used together with
Russian: NtechLab FaceNGN + VisionLabs Luna + Tevian + RecFaces, video tracking: ByteTrack (fastest) + DeepSORT + BoT-SORT + StrongSORT + MOTRv2. Video models: VideoMAE + X-CLIP + InternVideo, 3D vision: NeRF + Instant-NGP (NVIDIA) + Gaussian Splatting (rising 2023–2026) + Nerfstudio + Open3D + PyTorch3D (Meta) + Kaolin (NVIDIA) + COLMAP (SfM) + Mitsuba 3, generative: Stable Diffusion SDXL/SD3/SD3.5 + FLUX.1 (Black Forest Labs) + DALL-E 3 + Imagen 3 + Midjourney v6 + Ideogram. Video: Runway Gen-3 + Pika 1.5 + Kling + Sora + AnimateDiff + SVD.
Demand by grade
| grade | jobs |
|---|---|
| senior | 10 |
| junior | 3 |
| middle | 3 |
| lead | 1 |
Junior — typical entry: MS / PhD CV / Robotics + portfolio (Kaggle CV competitions / GitHub CV projects). Career flow: Backend Senior / ML Middle (2-3 years) + CV interest + portfolio → CV Engineer Junior (1-2 years) → Middle (2-3 years) → Senior → either 3D Vision Engineer (NeRF / Gaussian Splatting), Edge CV Engineer (TensorRT / CoreML mastery), Generative AI Engineer (image/video), ML Research CV (academic-track papers CVPR / ICCV / ECCV), or Autonomous Vehicle Perception Engineer. Numbers based on a small sample — for broader benchmarks see ml-engineer / research pages.
⚠ demand side only: the grade of the observed pool is unknown for most of it
Employers
Demand is spread across 21 employers. The largest accounts for 12.0%, the top ten for 56.0%; the remaining 44.0% is long tail.
| share | value |
|---|---|
| top-1 | 12.0% |
| top-3 | 28.0% |
| top-10 | 56.0% |
| long tail | 44.0% |
⚠ names are not shown: staffing agencies and end employers are not yet told apart by the classifier
Where Zorky sees this market
Observed across 14 sources; the largest accounts for 24.1% — this market does not rest on a single channel.
Recent openings
- Junior Computer Vision / ML Engineer · UA
- Senior I ML/CV Engineer · LT
- Machine Learning / Computer Vision
- Computer Vision Engineer · CH
- Computer Vision Engineer · MA
- Machine Learning / Computer Vision
- Senior / Lead Computer Vision Engineer · FR
- Senior Machine Learning Engineer (Robotics & Computer Vision) · PL
Latest open Computer Vision Engineer jobs — the most recent positions in the sample (narrow pool of explicit CV roles — the real market is wider thanks to overlap with ml-engineer / robotics). The full list is in our CRM or via the "see all" link below. For a broader view see ml-engineer + ai-engineer pages.
Adjacent markets
Computer Vision Engineer overlaps with ML Engineer (production ML overlap ~60%), AI Engineer (multimodal LLM overlap), NLP Engineer (vision-language models), Research Engineer (CVPR / ICCV / ECCV papers track), Robotics Engineer (SLAM + perception), Edge ML / MLOps Engineer (deployment overlap). Comparison with ml-engineer/ai-engineer/nlp/data-scientist/research/mlops — in the SiblingSubnichesChart above.
⚠ adjacent markets for comparison are not defined yet
How this is measured
- Vacancy
- an open job that cleared the quality gate and lists at least two technologies
- Observed candidate
- a candidate whose stack contains this technology; an aggregate — not a single record leaves the perimeter
- Matchable candidate
- not counted yet
- Window
- jobs open at the moment the snapshot was built
About the data
- Some breakdowns are hidden: their data coverage is not yet sufficient.
- Statistics are shown only where the sample clears a quality gate.
- A missing block does not mean a value of zero.
Breakdowns currently hidden: 9.
Data as of 2026-09-27
Direction: AI / ML / Data Science
Related specializations
Frequently asked questions
The most common questions about Computer Vision Engineer: pay (premium segment for rare-skill), CV vs ML vs AI Engineer (3-way comparison + 5 distinctions), object detection stack 2026 (YOLO vs Detectron2 vs MMDetection vs DETR vs SAM decision tree), 3D Vision Engineer (rising 2024+ sub-specialisation), remote, how to become (6-12 months from Backend / ML Middle + portfolio), Senior skills (PyTorch deep + OpenCV mastery + detection / segmentation frameworks + edge deployment + one domain). Answers recompute automatically.
What does a CV Engineer Junior, Middle, Senior, or Lead earn?
Numbers based on a small sample — for broader benchmarks see ML Engineer and Research Engineer / Scientist. Junior — typical entry: MS / PhD CV / Robotics + portfolio (Kaggle CV competitions, GitHub CV projects). The Junior → Middle jump — after the first production CV deployment (detection / segmentation / classification feature shipped). Middle → Senior — multi-task CV pipeline ownership + edge deployment expertise or generative AI specialisation. Senior → Staff / Principal — org-wide CV strategy + custom architecture design + research-paper publication track. Career flow: Backend Senior / ML Engineer Middle (2-3 years) + CV interest + portfolio → CV Engineer Junior (1-2 years) → Middle (2-3 years) → Senior → either 3D Vision Engineer, Edge CV Engineer, Generative AI Engineer (image/video), or ML Research CV (academic-track papers CVPR / ICCV / ECCV).
What stack does CV Engineer most often need?
Top 5: python, pytorch, c++, tensorflow, opencv. Python deep (mostly — edge components sometimes C++/Rust). PyTorch + torchvision mastery — 90%+ of CV research on PyTorch in 2026 (TensorFlow legacy). OpenCV mastery: classical CV ops (cv2.findContours / cv2.HoughLines / Canny edge / morphology / homography / camera calibration / stereo matching — must for preprocessing + traditional algorithms). Pillow (PIL — image manipulation). Object detection / segmentation: Ultralytics YOLO (YOLOv8/v9/v10/v11 — industry standard 2026 — easy training + good docs), Detectron2 (Meta — research-grade, more flexible), MMDetection / MMSegmentation / MMPose (OpenMMLab — huge ecosystem), DETR family (RT-DETR rising 2026), SAM + SAM 2 (universal segmentation). Self-supervised + foundation models: DINOv2 (Meta — best vision backbone 2024 for feature extraction), CLIP + OpenCLIP + SigLIP (vision-language alignment), MAE (Masked Autoencoder). Classification backbones: ViT family, Swin Transformer, ConvNeXt, EfficientNet + MobileNet (efficient), timm (pretrained models, must-library). Data augmentation: albumentations (industry standard — 30+ augmentation types), torchvision.transforms.v2 (modern PyTorch native), RandAugment + AugMix + Mixup + CutMix. Multimodal LLM: GPT-4V + Claude 3.5 Sonnet vision + Gemini 1.5/2.0 Pro vision + open-source LLaVA + Qwen 2 VL + InternVL + Idefics2 + Molmo (best open multimodal 2024). OCR: PaddleOCR (Baidu — multilingual leader 2026 — 80+ languages), EasyOCR, Tesseract (legacy still huge), Docling (IBM 2024 — doc understanding), Surya, TrOCR (Microsoft transformer-based). Face recognition: InsightFace (open-source SOTA 2026 — ArcFace + RetinaFace + Buffalo models), DeepFace, face_recognition. Russian: NtechLab FaceNGN / VisionLabs Luna / Tevian / RecFaces. Video understanding + tracking: ByteTrack (fastest 2026), DeepSORT, BoT-SORT, StrongSORT, MOTRv2. Video models: VideoMAE, X-CLIP, InternVideo, MViT. 3D vision: NeRF (original 2020) + Instant-NGP (NVIDIA — 100× faster) + Gaussian Splatting (rising 2023–2026), Nerfstudio (unified framework), Open3D (point cloud), PyTorch3D (Meta), Kaolin (NVIDIA), COLMAP (structure-from-motion), Mitsuba 3 (differentiable rendering). Generative image / video: Stable Diffusion family (SDXL + SD3 + SD3.5), FLUX.1 (Black Forest Labs — open record-holder 2024+), DALL-E 3 / Imagen 3 / Midjourney v6 / Ideogram. Video: Runway Gen-3 / Pika 1.5 / Kling AI / Sora / AnimateDiff / SVD. UIs for generative: ComfyUI (node-based — research favourite), Automatic1111. Fine-tuning generative: LoRA + DreamBooth for Stable Diffusion / FLUX. Inference serving: Triton Inference Server + BentoML + TorchServe. Annotation: CVAT (OpenCV — industry standard), Label Studio + V7 Darwin (commercial) + Roboflow (rapid prototyping). Datasets: COCO (detection / segmentation benchmark) + ImageNet (classification) + OpenImages + LAION-5B (web-scale generative training) + Objects365 + Visual Genome + ADE20K (segmentation) + Cityscapes (autonomous driving). Hardware-aware optimisation: quantisation (INT8 / FP16), pruning, distillation, structured sparsity (for NVIDIA Ampere+ tensor cores).
Computer Vision Engineer vs ML Engineer vs AI Engineer — what's the difference?
All three roles overlap in 2026, but focus areas differ. ML Engineer — generalist, owns the production ML stack (recsys / fraud / ranking / classical ML + LLM). Can work with CV data but no deep specialisation. See ML Engineer. AI Engineer / LLM Engineer — focus on LLM integration (text-focused). See AI / LLM Engineer. NLP Engineer — focus on text + speech. See NLP Engineer. Computer Vision Engineer (this page) — focus on image / video / 3D data. Stack overlap with ML Engineer ~60% (PyTorch + cloud ML + deployment) + ~30% unique (OpenCV + Ultralytics + Detectron2 + SAM + albumentations + edge deployment specifics — TensorRT / CoreML). Distinctions: 1) Image-processing intuition — CV Engineer understands camera models / lens distortion / colour spaces / histogram analysis / morphological ops — classical CV foundation. AI Engineer usually knows none of this. 2) Computational constraints — CV models are heavy (gigabytes), inference latency critical (autonomous driving — 30+ FPS mandate), edge deployment (TensorRT / CoreML / OpenVINO mastery) — exclusive CV territory. 3) Generative image / video specialisation — Stable Diffusion / FLUX / ComfyUI workflow expertise — CV-niche specifically (overlaps with AI Engineer for text-driven generation). 4) 3D vision skills — NeRF / Gaussian Splatting / point cloud / SLAM / camera calibration — exclusive CV. 5) Domain expertise — CV roles often tied to a specific industry (autonomous vehicles / medical imaging / robotics / AR/VR / satellite imagery / manufacturing QC). Each domain has its own datasets + rules. Career pivots: ML Engineer Senior → CV Engineer — 4-8 months (need OpenCV + Ultralytics + Detectron2 + albumentations + one domain). CV Engineer Senior → ML Engineer — 2-4 months (easy lateral, add MLOps + LLM basics). CV Engineer Senior → AI Engineer (LLM track) — 3-6 months. Hot 2025–2026 sub-specialisations: 3D Vision (Gaussian Splatting) / Generative Image-Video (FLUX + Sora-style) / Multimodal (Vision-Language Models — LLaVA / Molmo).
Object detection stack 2026 — YOLO vs Detectron2 vs MMDetection vs DETR vs SAM?
Decision tree for object detection / segmentation 2026: 1) Ultralytics YOLO (YOLOv8 / v9 / v10 / v11) — default choice 2026 for object detection. Pros: easy training (3-5 lines of code), good docs + community, fast inference (real-time on CPU + edge), wide deployment support (ONNX + TensorRT + CoreML + TFLite native). Cons: Ultralytics license requires AGPL or commercial license ($$) for proprietary use, less academic-flexible than Detectron2 / MMDetection. Use case: 90% of production object detection use cases, prototypes, edge deployment. 2) Detectron2 (Meta) — research-grade. Pros: clean API, flexible (easy custom architectures), Apache 2.0 license (commercial-friendly), backed by Meta. Cons: slower training than Ultralytics, larger learning curve. Use case: complex custom architectures, research projects, when you need fine control over the training loop. 3) MMDetection / MMSegmentation / MMPose (OpenMMLab — Chinese consortium) — huge ecosystem. Pros: 100+ pre-implemented architectures (old and new), papers' reference implementations always there first, Apache 2.0. Cons: config-heavy (steep learning curve), Chinese-language community sometimes hard for English-only, dependencies can conflict. Use case: research, comparing many architectures, paper reproduction. 4) DETR family (transformer-based detection) — modern paradigm shift 2020+. RT-DETR (Real-Time DETR — Baidu, rising 2024+ — competitor to YOLO in real-time space), DINO-DETR, Co-DETR. Pros: end-to-end (no NMS post-processing), often better accuracy on complex scenes. Cons: slower inference than YOLO, more compute-hungry for training. Use case: highest-accuracy needs, complex scenes (crowded objects). 5) SAM (Segment Anything Model — Meta 2023) + SAM 2 (video version 2024) — universal segmentation. Use case: a) zero-shot segmentation (segment anything without training), b) interactive annotation tool (point / box prompt → mask), c) data-labelling acceleration (SAM-assisted annotation 10× faster). Combine with YOLO / Detectron2: detector → SAM for precise masks. 6) Classical detection (HOG + cascade classifiers — OpenCV) — only for very simple cases, edge devices without GPU, low-power microcontrollers (still relevant for embedded scenarios). Default 2026 recommendations: Production deployment + commercial → Ultralytics YOLO (if OK with AGPL or commercial license) OR Detectron2 (Apache 2.0 alternative). Research / paper reproduction → MMDetection. Highest accuracy / complex scenes → RT-DETR or Co-DETR. Universal / zero-shot segmentation → SAM 2. Annotation acceleration → SAM + interactive workflows (CVAT integration). Edge / mobile → YOLO with TensorRT (NVIDIA) / CoreML (Apple) / TFLite (Android). A Senior CV Engineer must know when to use which.
Can CV Engineers work remotely?
Yes, 24% of CV Engineer jobs are full-remote or hybrid. CV work is fully cloud-based (training in cloud GPUs — A100 / H100, datasets streaming from S3 / GCS, deployment in Kubernetes). Outsourcing shops — almost always remote on US CV projects. Russian — hybrid or remote after probation. Russian banks (Sber AI Banking CV — face recognition / document verification) — hybrid/office security compliance. Autonomous vehicle companies — special case: research roles often remote-friendly, but product deployment / on-vehicle testing requires on-site (Pittsburgh Cruise / Mountain View Waymo / Palo Alto Tesla / etc). Generative AI companies (Stability / Black Forest Labs / Runway / Pika / Kling) — full-remote standard. International voice-AI / NLP companies overlap (multimodal teams) — full-remote. Big Tech CV — hybrid-standard. Relocant hubs for CV: USA (Bay Area + Pittsburgh — autonomous vehicle clusters + Boston — Robotics MIT region), UK (London — Wayve), Canada (Toronto — Vector Institute), Germany (Berlin AI + Munich automotive — BMW / Mercedes AI), France (Paris — Hugging Face + Mistral for vision LLM), Japan (Tokyo — Sony AI + automotive). English for international CV remote — must (CVPR / ICCV / ECCV papers + community English-speaking).
How is 3D Vision Engineer (NeRF / Gaussian Splatting — rising 2024+) different?
3D Vision Engineer — sub-specialisation within CV focused on 3D understanding + neural rendering. Day-to-day: 1) NeRF / Gaussian Splatting reconstruction — input video → 3D scene representation. Tools: Instant-NGP (NVIDIA — fast NeRF) / Nerfstudio (unified framework) / gsplat (Gaussian Splatting library) / Polycam app for capture. 2) Point cloud processing — LiDAR data (autonomous vehicles) or depth cameras (Kinect / RealSense). Tools: Open3D / PCL (Point Cloud Library) / PyTorch3D. 3) Structure-from-motion (SfM) — multiple 2D images → 3D scene. Tools: COLMAP (industry-standard SfM). 4) SLAM (Simultaneous Localization and Mapping) — robotics / AR — track camera pose + build map. Tools: ORB-SLAM3 / OpenVSLAM / Kimera. 5) Differentiable rendering — learn 3D from 2D supervision. Tools: Mitsuba 3 / nvdiffrast / Kaolin. 6) 3D generation — text → 3D mesh (DreamFusion + Magic3D + Zero-1-to-3) or image → 3D (TripoSR / InstantMesh — open-source 2024+). 7) Mesh processing — texturing / retopology / UV unwrapping for 3D content pipelines. Stack-specific: PyTorch3D (Meta — 3D research), Kaolin (NVIDIA — 3D DL), Mitsuba 3 (differentiable rendering), Open3D / PCL (point cloud), COLMAP (SfM), Nerfstudio + gsplat (NeRF / 3DGS), Blender Python API (mesh manipulation). Career flow: CV Engineer Senior + computer graphics interest + NeRF / 3DGS hands-on portfolio → 3D Vision Engineer — 6-12 months.
Where to start in Computer Vision in 2026?
Roadmap: 1) Math foundations — linear algebra + calculus + basics of projective geometry (transformations / homographies / camera models). "Multiple View Geometry in Computer Vision" Hartley / Zisserman — bible for CV math (can be used as a reference, no need to read end-to-end). 2) Python deep + ML basics — PyTorch + NumPy + Matplotlib. Build a simple ML classifier (MNIST). 3) OpenCV mastery — classical CV ops. Course: "Computer Vision Course" by PyImageSearch (Adrian Rosebrock — best entry for OpenCV), OpenCV official documentation + tutorials. Build pet projects: edge detection / homography panorama stitching / face detection with classical Haar cascades. 4) Deep Learning for Vision — Stanford CS231n "Convolutional Neural Networks for Visual Recognition" (Karpathy / Li — free YouTube + slides — must-do, canonical CV deep-learning course). 5) PyTorch + torchvision hands-on — train ResNet / EfficientNet on CIFAR-10, fine-tune a pretrained model on your own dataset. Learn to use timm library (pretrained models). 6) Object detection — Ultralytics YOLO hands-on (easiest entry — train YOLO on a custom dataset in a day). Then Detectron2 (more flexible). 7) Segmentation — Mask R-CNN via Detectron2, then SAM (universal segmentation) — try interactive segmentation. 8) Annotation tooling — CVAT mastery (industry standard for CV annotation). Annotate your own dataset (10-50 images), train detector on it. 9) Augmentation mastery — albumentations library (must for production CV). Understand training stability — strong augmentation improves robustness. 10) Modern transformers for CV — Vision Transformer (ViT) + Swin Transformer + DINOv2 (foundation model). Hugging Face Transformers vision support. 11) Multimodal LLM — try GPT-4V + Claude vision + open-source LLaVA / Qwen 2 VL for understanding. 12) Generative CV track (popular 2024–2026): Stable Diffusion + FLUX hands-on, ComfyUI workflows mastery, LoRA / DreamBooth fine-tuning. Course: "Generative AI with Diffusion Models" DeepLearning.AI. 13) 3D Vision track (rising 2024+): NeRF + Gaussian Splatting hands-on with Nerfstudio framework, capture your own scene with a phone (Polycam app), reconstruct as 3DGS. 14) Edge deployment hands-on — convert PyTorch model → ONNX → TensorRT (on NVIDIA GPU), benchmark inference latency. Try CoreML conversion for Apple Silicon. 15) Pet project portfolio: a) production-grade detection pipeline (e.g. fish-counting / vehicle-counting / cell-counting demo); b) custom Stable Diffusion / FLUX LoRA on own style; c) Gaussian Splatting reconstruction (cool 3D scene from phone video); d) mobile CV app (deploy via CoreML/TFLite). Document on GitHub + blog post + video demo. Russian courses: MIPT DLSchool (CV module — free YouTube), Karpov.Courses "Computer Vision" track, Otus "Computer Vision", SkillFactory CV, School21 (Sber) AI Computer Vision track. International (EN): Stanford CS231n (canonical — free YouTube), fast.ai Practical Deep Learning (CV included), Hugging Face Computer Vision Course (free), "Deep Learning for Computer Vision" book Mohamed Elgendy, PyImageSearch University (Adrian Rosebrock — applied focus). Must-read books: "Deep Learning for Vision Systems" Mohamed Elgendy (Manning), "Multiple View Geometry" Hartley / Zisserman (math reference), "Computer Vision: Algorithms and Applications" Richard Szeliski (free 2nd edition online — encyclopaedic). Conferences: CVPR (top — June), ICCV (October, alternating with ECCV), ECCV (October alternating), NeurIPS (December — broader ML with CV track). Backend Senior / ML Engineer Middle + CV interest + portfolio → CV Engineer Junior — 6-12 months. PhD CV / Robotics → Senior CV Engineer — direct entry.
How many CV Engineer jobs are open across CIS and Europe?
29 active open CV Engineer positions with explicit CV specifics in our sample. The real pool is many times wider — many CV roles are classified as general ML Engineer / Robotics / AI Engineer (titles like "ML Engineer for autonomous driving" or "Senior Backend Engineer with CV focus"). True CV-focused jobs in CIS + Europe are estimated at 300-1,500 positions active at any moment in 2026 (counting fuzzily classified ones). Geography: Russia / Poland / remote. The real market is broader thanks to the international remote segment (generative AI companies + robotics startups full-remote-friendly). Time to close a Senior CV Engineer — 6-12 weeks (longer than general AI Engineer due to rare-skill — PyTorch deep + classical CV + one domain + edge deployment combination).
What skills does a Senior CV Engineer need?
A Senior CV Engineer owns the full vision-engineering cycle + technical leadership. Math foundations: linear algebra + projective geometry (camera models / homographies / epipolar geometry) + calculus + optimisation theory — at the level of "can read CVPR papers without math blocks". Python deep + Backend Senior level: async / typing / FastAPI / pytest mastery. C++ basics (for edge deployment + OpenCV custom kernels — nice to have). PyTorch + torchvision mastery deep: custom Datasets / Samplers / Losses / training loops, distributed training (DDP for multi-GPU CV training), mixed-precision (FP16 / BF16 mandatory for production), gradient accumulation for large models. OpenCV mastery: classical CV ops (calibration / morphology / Hough transforms / homography / stereo matching / optical flow), C++ API basics for performance-critical paths. Modern transformers for CV: ViT / Swin / DINOv2 mastery, fine-tuning strategies, hybrid CNN-transformer architectures. Object detection / segmentation mastery: Ultralytics YOLO production deployment + Detectron2 custom architecture authoring + MMDetection / MMSegmentation when needed + RT-DETR + SAM 2 integration in pipelines. Data augmentation mastery: albumentations advanced (custom transforms + composition strategies), test-time augmentation (TTA), AutoAugment / RandAugment / CutMix / Mixup understanding. Foundation models: DINOv2 + CLIP + SigLIP + OpenCLIP — use cases for feature extraction + zero-shot classification + retrieval. Multimodal LLM (GPT-4V / Claude vision / LLaVA / Qwen 2 VL / Molmo) — when to use vs train custom model. Generative CV mastery (if track includes): Stable Diffusion + FLUX deep — ControlNet / LoRA / DreamBooth / IP-Adapter / textual inversion fine-tuning. ComfyUI workflow authoring. Video generation basics (AnimateDiff / SVD). 3D Vision mastery (if track includes): NeRF + Gaussian Splatting + Nerfstudio + COLMAP SfM + Open3D point clouds. Camera calibration + projective geometry deep. Edge deployment mastery: ONNX (cross-framework export + ONNX Runtime optimisation), TensorRT (NVIDIA — INT8 quantisation + plugin development if needed), CoreML (Apple Silicon — Conv2D ops support + Neural Engine specifics), TFLite (Android NNAPI delegate), MediaPipe pipelines, OpenVINO (Intel iGPU). Hardware-aware optimisation: quantisation (PTQ + QAT), pruning, distillation (teacher-student), structured sparsity (NVIDIA Ampere+ for 2× speedup). Production CV pipeline architecture: design end-to-end pipeline on a whiteboard — data ingestion (cameras / drives / cloud streaming) → preprocessing → batched inference → post-processing → tracking / aggregation → downstream actions. Multi-camera systems: synchronisation + calibration + fusion for autonomous / surveillance. Inference serving: Triton Inference Server (NVIDIA — best for multi-model + dynamic batching), DeepStream (NVIDIA video analytics), BentoML / TorchServe. Latency budgets: design for real-time constraints (autonomous: 30+ FPS, AR: 60+ FPS, video moderation: batch ok), profile-driven optimisation (NVIDIA Nsight + PyTorch Profiler). Domain expertise: deep understanding of one or two domains (autonomous driving / medical imaging / retail visual search / robotics perception / AR/VR / satellite imagery / manufacturing QC) — the main premium driver for Senior+. Annotation strategy: design large-scale annotation workflows (CVAT / Label Studio + SAM-assisted), active learning loops, label quality control. Evaluation methodology: COCO metrics (AP / AP50 / AP75 / mAP@[.5:.95]) deep understanding, segmentation metrics (mIoU / Dice / boundary F1), domain-specific metrics (NDS / mAP for autonomous detection), human-in-the-loop evaluation for generative. Soft: ADRs writing for CV architecture decisions, technical writing (CV feature design docs + paper drafts if research-track), cross-team collaboration (Product / Backend / Robotics / Hardware teams), mentoring Middle CV Engineers, paper-reading discipline (CVPR / ICCV / ECCV / NeurIPS must-follow). English for Senior+ MUST — CV community / docs / papers / conferences are English-speaking. Optional bonus: open-source contributions to Ultralytics / Detectron2 / MMDetection / timm / albumentations / Nerfstudio — sharply increase market value for Big Tech CV / autonomous vehicles / generative AI hiring. Papers at CVPR / ICCV / ECCV workshops — premium for Research Engineer CV track.
Leave a request
Describe the task and leave a contact — the request goes to our CRM and we reply at the contact you provide.