Service · Software development

Custom computer vision application development.

Custom computer vision that really works in your production environment, warehouse, shop or field. From object detection and image recognition to OCR, segmentation and anomaly detection, built as an integration into your existing software, not as a standalone proof of concept.

Object detectionImage recognitionOCREdge & cloud

Computer vision is AI on images, and it has been running in production for years.

Cameras can now see more than people can keep track of. On an assembly line, more product passes by in an hour than a QA inspector can visually check. In a warehouse, thousands of labels are scanned. On a construction site, someone always has to check whether helmets are being worn. Computer vision (CV) automates these kinds of tasks — not as magic, but as a reliable part of your operational software.

We build computer vision applications that slot straight into your existing stack: a YOLO model running on an edge device next to the production line, an OCR pipeline that reads invoices and posts the lines into your accounts, or a classification model that pre-filters insurance claims for your back office. Not a standalone "AI pilot" that never goes live, but a working integration with your ERP, WMS or CRM. For strategic planning, we also help with a broader AI strategy in which computer vision is one of the applications.

In recent years the tooling has matured quickly. Pre-trained models such as YOLO, Detectron2, Segment Anything and Vision Transformers have taken away much of the heavy lifting: you no longer need to label millions of images, but a few hundred to a few thousand per class. At the same time, edge devices have become powerful enough to run those same models locally — next to the production line, in a drone or in a barrier installation. And multimodal LLMs such as GPT-4o-vision, Claude Vision and Gemini Vision give you a new, shorter route for zero-shot and few-shot tasks for which you would previously have trained your own model. Which route fits depends on your data, your latency requirements, your privacy context and the business case around it — and that is exactly what the first conversation is about.

Three types of computer vision projects.

Depending on what you need to recognise, how critical it is, and whether existing models are good enough. In the first conversation we advise which type fits.

Compact project · fixed sprint budget

Cloud API integration with an existing model

For standard tasks — general object detection, OCR on regular documents, face blurring for GDPR compliance — we use Google Vision, Azure Computer Vision, AWS Rekognition or an open-source pre-trained model. Fast implementation, no proprietary training set required, pay per call.

Google VisionAzure CVOCRAPI integration
Mid-sized project · fixed sprint budget

Custom model on your own data

When you need to recognise something specific — your own products, defects nobody else spots, industry-specific documents — we fine-tune a pre-trained model (YOLOv8, Detectron2, Vision Transformer) on your imagery. Includes the labelling process, training pipeline and inference service.

YOLOv8/v10Fine-tuningLabel StudioPyTorch
Larger project · fixed sprint budget

Edge deployment or real-time video pipeline

For applications with low-latency requirements — live production inspection, perimeter surveillance, vehicle tracking — the model runs on an edge device (NVIDIA Jetson, Google Coral, Hailo) or within a real-time video pipeline. We build the entire stack: capture, pre-processing, inference, post-processing and dashboard.

NVIDIA JetsonCoreML / TFLiteByteTrackReal-time video

What you get at the end.

A working CV system you can maintain yourself, plus everything around it to build further.

  • The trained model + inference serviceA production-ready model (PyTorch, ONNX or CoreML) plus the service that feeds it images and returns results via REST or message queue.
  • Integration with your softwareIntegration with your existing ERP, WMS, MES, CRM or camera system, so that detected objects, OCR output or anomaly flags land directly in the right workflow.
  • Training data + labelsThe labelled dataset in an open format (COCO, YOLO, Pascal VOC) so you can retrain it yourself or hand it over to another team.
  • Training pipeline and retraining protocolA reproducible pipeline (Weights & Biases or MLflow) so you can retrain when the model drifts, without having to call us.
  • Monitoring and evaluationDashboards for precision, recall and confusion matrices in production, plus drift alerts when the input distribution changes.
  • Documentation and maintenance contract (optional)Architecture overview, GDPR and AI Act documentation where relevant, and ongoing maintenance for security patches and model updates.

Use cases we build in practice.

Computer vision touches many sectors. Below are the scenarios where we support clients. If you recognise one of them, we would be happy to talk further.

Production

Quality and defect detection

Cameras above the assembly line detect scratches, dents, missing components or incorrect colours. A YOLO or anomaly detection model (PatchCore, PaDiM) flags defective units so they are sorted out automatically before your QA team even sees them.

Logistics

Label OCR and stock counting

Parcels and pallets moving through your warehouse, labels and barcodes on damaged boxes, and automatic counting of stock on shelving. We combine OCR (PaddleOCR, Google Vision) with object detection and integrate it with your WMS.

Insurance & property

Photo-based damage assessment

A claimant uploads a photo, a classification model assesses the type and severity of the damage, and the claim is pre-sorted. The same pattern works for property inspections, drone photos and pre-acceptance. It fits well with custom LLM integrations, where text and images are analysed together.

Retail & agriculture

Shelf monitoring and crop analysis

Shelf cameras recognise out-of-stock situations and planogram deviations. In agriculture we do similar work with drone imagery: crop stress, fruit ripeness and livestock counts in a field, using segmentation on satellite or drone photos.

Construction & safety

PPE and incident detection

Detection of helmets, high-visibility vests and safety goggles on construction sites, combined with perimeter monitoring. We only work on applications that are ethically and legally sound: no biometric mass surveillance, no emotion recognition, but safety within defined contexts.

Mobility

Licence plate and vehicle recognition

ANPR for car parks, charging points or barriers. Vehicle tracking for traffic analysis and occupancy. Licence plate data is personal data under the GDPR, so we build in appropriate retention and storage periods.

Healthcare

Medical image analysis

For applications involving medical imaging, we work only via an MDR-certified route, often together with a medical device partner. For non-diagnostic tasks (workflow support, anonymisation, image routing) we build in-house to NEN 7510 requirements.

OCR & documents

Invoice, passport and form OCR

Extracting text and structured data from photos or scans, with validation against your master data. We combine PaddleOCR or Tesseract with multimodal LLMs for the difficult edge cases where rule-based OCR falls short.

Not yet sure about a large project?

Test your idea first: a working prototype in 1 day

With OneDayBuild, we turn your idea into something tangible in one day for €1,150, so you can see whether further development is worth the investment. Decide to go ahead with the full build? Then we credit the full cost.

Explore OneDayBuild →

Tech stack we work with.

A brief tour of the models, frameworks and deployment targets we use in practice. We choose per project; there is no one-size-fits-all stack.

Models

Detection, segmentation and classification

YOLOv8, YOLOv9 and YOLOv10 for real-time detection. Detectron2 and MMDetection for heavier tasks. Segment Anything (SAM) for pixel-precise boundaries. Vision Transformers (ViT, DINOv2) where the transformer architecture offers an advantage.

Multimodal LLMs

GPT-4o, Claude Vision, Gemini Vision

For zero-shot tasks, complex reasoning on images, or where text and images need to be analysed together. Often a quick way to validate whether CV adds any value at all before we train a custom model.

Frameworks & tooling

PyTorch, ONNX, OpenCV

PyTorch as the training framework, ONNX for cross-platform export, OpenCV for pre- and post-processing. Tracking with SORT, ByteTrack or DeepSORT. Anomaly detection with PatchCore or PaDiM for industrial inspection.

Deployment

Cloud, edge and on-device

Cloud GPUs for batch processing and training, NVIDIA Jetson and Google Coral for edge, CoreML and TFLite for mobile and on-device. Inference is packaged as a REST, gRPC or WebSocket service, suitable for both real-time and batch use.

Data labelling & MLOps

Label Studio, CVAT, V7

For labelling we work with Label Studio, CVAT or V7. We track experiments and model versions in Weights & Biases or MLflow. CI/CD pipelines for training and deployment, plus monitoring of precision, recall and drift.

Synthetic data

Stable Diffusion and 3D renderers

When you have too few production images, we generate synthetic data: Stable Diffusion for variation, or a 3D renderer for pixel-precise labels. This is often the key difference between "it works in the demo" and "it works in production".

How a computer vision project runs.

1

Use case and feasibility

A conversation in which we sharpen the use case: what needs to be recognised, how often, how quickly, and with what accuracy it becomes valuable. We often carry out a short feasibility scan to determine whether an existing cloud model is sufficient or whether training is needed.

2

Data collection and labelling

Almost every CV project stands or falls with the data. We help collect representative images, choose the right labelling platform (Label Studio, CVAT, V7) and agree an acceptance criterion for dataset quality.

3

Model training and evaluation

We fine-tune a pre-trained backbone on your data, track experiments in Weights & Biases or MLflow and validate against a hold-out set. You receive confusion matrices, precision/recall charts and concrete examples of what the model does and does not see, so you know what you are buying.

4

Integration and deployment

The model is packaged as an inference service (REST, gRPC or a converted CoreML/TFLite file for edge) and integrated with your existing software. We work in sprints and deliver something you can test every two weeks, with monitoring built in from day one.

5

Production and ongoing development

After go-live, we monitor for model drift, retrain periodically on new data and, if you wish, offer a maintenance contract under which we evolve the model alongside your operations.

Frequently asked questions.

What clients usually want to know before we start.

How can I integrate my software development with computer vision?
Almost always through a separate inference service that your existing application calls. Your software sends an image (or streams frames), and the service returns a structured response with objects, bounding boxes, confidence scores or OCR text. For real-time scenarios we use message queues or WebRTC; for batch processing a simple REST API works well. The computer vision component thus becomes a normal subsystem in your architecture, not something exotic. In practice we already connect it to ERP, WMS, CRM and MES systems, and we can also work within a broader AI agent architecture in which several models work together.
What are the most common object detection applications?
Object detection finds objects in an image along with their position (bounding box). Concrete applications we often build include: defect detection on production lines, package and pallet detection in warehouses, helmet and PPE detection on construction sites, shelf monitoring in retail, vehicle and number-plate detection, and animal and plant detection for agriculture. We typically work with YOLOv8 or YOLOv10 for speed, Detectron2 for heavier tasks, and DETR variants where the transformer architecture offers an advantage.
What is the difference between image recognition, classification and object detection?
Image recognition is the umbrella term. Classification assigns a single label to a whole image ("this is a defective part"). Object detection finds multiple objects within an image and gives their location ("two scratches at this position, one dent there"). Segmentation goes a step further and provides pixel-precise boundaries. Which variant you need depends on the use case: for "is this broken or not" classification is sufficient, while for "where is the defect" you need object detection.
Cloud API or a custom-trained model?
For standard tasks — general OCR, face blurring, generic object detection — cloud APIs from Google, Azure or AWS work well and you pay per call. For specific domains (your own products, your own defects, your own documents), generic cloud models often perform too poorly and a custom fine-tuned model is needed. We usually start with a cloud API to validate the business case, and only move to a custom solution when the figures call for it.
How much training data do I need?
With pre-trained models, a few hundred to a few thousand labelled examples per class is often sufficient. For clearly distinguishable classes, even fewer; for subtle differences or many classes, more. Augmentation and synthetic data (generated with Stable Diffusion or a 3D renderer) can help to supplement a smaller dataset. In the feasibility conversation we give a reasoned estimate for your specific problem.
Edge or cloud — where does the model run?
Cloud is simple, scalable and inexpensive to maintain, but it adds latency and sends images over the network. Edge inference on NVIDIA Jetson, Google Coral or Hailo is faster, works offline, and keeps images local — important where privacy or bandwidth plays a role. For real-time production inspection we usually choose edge, for batch analysis cloud. A hybrid setup is also possible: edge for inference, cloud for monitoring and retraining.
What about the AI Act and the GDPR for computer vision?
Computer vision applications that identify or categorise individuals are subject to strict rules. Real-time biometric identification in public spaces, emotion recognition in workplaces or schools, and predictive policing are either prohibited or classified as high-risk under the AI Act. Faces and number plates are personal data under the GDPR and require a legal basis, a retention policy and often a DPIA. We do not undertake biometric identification or emotion recognition without a thorough legal process — for PPE detection, defect inspection or stock counting, this barely plays a role, if at all.
What determines the cost of a computer vision project?
The largest cost items are data labelling, model training (especially GPU time for larger models) and integration with your existing systems. A cloud API integration for a standardised task is a compact project; an edge deployment with custom training and a real-time video pipeline is a project spanning several sprints. We work with fixed sprint budgets so that you can reassess after each sprint.
Do you only work on ethically responsible computer vision?
Yes. No mass surveillance, no emotion recognition, no predictive policing, no facial identification without an explicit legal basis and DPIA. We do, however, take on quality inspection, stock counting, OCR, PPE detection, crop analysis, damage assessment for insurance, and similar functional applications. When in doubt, we carry out an AI Act assessment beforehand — if it fits within the framework, we build; if it doesn't, we'll tell you honestly. See also our page on enterprise AI implementation, where the same governance approach recurs.

Talk to us about your computer vision application.

A half-hour introductory call, no obligation. We listen to the use case, ask about your data and give direction on feasibility, approach and which models fit. If it sits within a broader AI landscape, we'll extend the conversation to generative AI solutions or a wider strategy.

Edit content