Edge Vision-Language Model (SmolVLM)

Project information

  • Category: MLOps
  • Client: INSTAR Deutschland GmbH (Waletech) Shenzhen, China
  • Project date: 15 Jul, 2026
  • Project URL: INSTAR Cloud

The problem: a VLM is a GPU-shaped tool, and we had a CPU server and a hard latency budget.

What I did: I fine-tuned SmolVLM on a custom security-scene dataset. To teach it what matters, I first ran my YOLO detector over the footage and fed those bounding-box regions into a Gemma-4 pass — using the detector's output to point the model at the regions of interest and prime it to be an expert on spotting potential security risks, before it ever sees the whole scene. I trained multiple variants (256M / 500M, 8-bit, different prompts, different output shapes) specifically to cut token-generation latency before it reached the serving layer: deterministic decoding, a strict 512-token budget, and enforced-JSON outputs. I exported the models to OpenVINO (OVModelForVisualCausalLM) for CPU deployment, and the serving layer keeps each variant loaded once behind a thread-safe lock so a worker thread returns a prediction without blocking the pipeline.

In production it runs inside the Nomad-managed inference job: model weights live on a shared host volume — a local model registry the container mounts — so shipping a new variant is a canary-safe container rollout, not a retrain-and-redeploy cycle.

In progress: feeding the scene descriptions into Elasticsearch so an alarm event becomes customer-searchable text.

The payoff: multimodal understanding on hardware that shouldn't be able to do it, within a performance target.

Designed with BootstrapMade