INSTAR Cloud - Vision Language Annotation

Project information

  • Category: MLOps
  • Client: Waletech (INSTAR Germany) Shenzhen, China
  • Project date: 15 Jul, 2020
  • Project URL: INSTAR Cloud

SmolVLM exported to the OpenVino format - finally a VLM transformer model that can be deployed on a CPU server! The model was fine-tuned using custom annotation using a YOLO object detector with a GEMMA4 VLM - letting the custom trained object detector point out points of interest in the custom dataset to the Vison-Language-Model. Different SmolVLM models were trained with different input prompts and output formats to minimize the output generation time. Still WiP - adding the predicted output to Elasticsearch to make alarm events search-able.

Designed with BootstrapMade