Face Recognition – PCA / HoG / SVM Pipeline

Project information

  • Category: MLOps
  • Client: INSTAR Deutschland GmbH (Waletech) Shenzhen, China
  • Project date: 01 Feb, 2025
  • Project URL: INSTAR Cloud

The problem: a heavy TensorFlow embedding model slowing the pipeline down, and a customer who wanted their custom face detections without filing a support ticket. What I did: Replaced the embedding model with a PCA projection and a Histogram-of-Oriented-Gradients fallback, while keeping a CNN-SVM (FaceNet) path for higher accuracy. The model swap was the easy half — the real build was the multi-tenant workflow: a self-serve pipeline the customer triggers from the cloud UI, where YOLO face crops get assigned to user-defined classes, quality-scored, and balanced (including an "unknown-faces" class so a one-class SVM rejects strangers), then flow through either the PCA-SVM or HoG-SVM classifier (CNN-SVM fallback for safety). Per-user models, label encoders, and face galleries are persisted on shared host volumes — a local model registry the inference container mounts — so any worker in the fleet can score a face against the same tenant's gallery without a database round-trip per request. I also kept parallel `dev` and `prod` model/gallery directories in the serving job, so a customer can A/B a retrained face list against the live one before it goes real. At inference time, face embeddings are clustered and grouped into a coherent list of persons, with threshold + margin scoring to avoid false matches. Shipped as a self-serve "train your own face list" feature. The payoff: moved face-recognition training from an internal task into the customer's own hands while cutting prediction times by up to 50% to sub 300ms.

Designed with BootstrapMade