OpenMV
Shields
Lens
Robotics
Book
Video
Download
Docs
Forum
OpenMV.io
GithubArtificial intelligence has two very common tasks in the field of vision: classification and object detection.
Category:Determine which category the image belongs to, for example, whether it is a cat or a dog.
Object Detection:Used to locate the positions, quantities, and dimensions of multiple distinct objects in an image.
In previous video tutorials, we explained how to use the OpenMV4 Plus to train neural networks online for classifying different objects. This enables real-time determination of whether an image within the OpenMV field of view corresponds to a specific object—for example, whether a face is wearing a mask—but it cannot output the position coordinates of faces or masks, nor the number of faces.
Target detection is significantly more complex and computationally intensive than classification, making it rare in the microcontroller domain. OpenMV’s main control MCU is the STM32H7; we have now added a new feature to SingTown that enables target-detection-like tasks to run on microcontrollers, delivering outstanding performance.

Today we will explainTraining demonstrates traffic sign detection in real-road environments, detecting traffic signs on our actual roadssuch as no-horn zones, no-parking zones, and speed limits of 80 km/h, rather than laboratory environments free from interference.
You can view the entire project content, including all code and models, at the following address: https://book.openmv.cc/project/traffic-sign.html
It is widely known that real-world environments—particularly road environments—are highly complex and subject to numerous sources of interference. We captured images of traffic signs in real-world environments using devices such as smartphones, then trained a model on the Edge Impulse online platform—a partner of SingTown’s OpenMV—by manually labelling approximately 270 images; the training process required only ten minutes. The resulting model achieved an F1 score accuracy of 92% and delivered a frame rate of ten frames per second on OpenMV, yielding both smooth performance and high accuracy.
Users can train the system to detect any target of interest, such as different numerals, various fruits, distinct markers, different components, or even any specific irregular object, using our video tutorials; this enables detection of the quantity, coordinates, and object class name of the specified target.
Note:This video tutorial covers the neural network keypoint detection feature, which is compatible not only with the OpenMV4 H7 Plus but also with the OpenMV4 H7. The trained FOMO keypoint detection model is compact, highly effective, and only tens of kilobytes in size; therefore, even though the OpenMV4 H7 has less RAM than the Plus variant, we can train a slightly smaller model to run on the OpenMV4.
We outline the underlying principle below:
The primary design decision behind FOMO object point detection on OpenMV is based on the idea that many object detection problems do not actually require the size of objects, but only their positions within an image. Once we know the positions of objects, subsequent operations can be performed—for example, using OpenMV to control a vehicle to move toward a detected target, enabling a drone to land at a specific location, or controlling a robotic arm to grasp a particular target object.
The FOMO model is a model specially designed by EdgeImpulse for microcontroller environments; unlike conventional object detection, it can detect only the positions and number of multiple objects, not their specific dimensions.
The underlying principle of the FOMO model is very simple and flexible. It first divides an image into patches, each measuring 8×8 pixels. For an image with a resolution of 96×96, this results in a grid of 12×12 cells. For an image with a resolution of 360×360, this results in a grid of 45×45 cells. Image classification is then performed independently on each cell.
FOMO performs significantly better than YOLO V5 or MobileNet SSD on large numbers of small objects. FOMO object point detection is 30 times faster than MobileNet SSD.
FOMO performs better when the objects to be detected are of similar size—for example, when the sizes of the markers to be identified are broadly consistent, rather than varying significantly between very large and very small objects.
OpenMV uses EdgeImpulse to train neural network object detection models online, which mainly involves the following steps: collecting and uploading the image dataset, labelling objects, training, testing the model, and deployment.
* Collecting an image dataset: Capturing images of specific objects in the actual environment where recognition will be performed. Images are collected in the same environment where recognition will take place, ensuring that the actual training environment closely matches the actual detection environment, thereby achieving better results.
* Annotation Target: Draw bounding boxes around each target to be detected in our dataset and label them with the corresponding target names to facilitate subsequent training for detecting both target locations and names.
* Training the model: After annotation is complete, train the model using the convolutional neural network parameters designed by us.
* Test model: We conduct testing using the trained model; if the performance is unsatisfactory, we correspondingly expand the training dataset or adjust the training parameters and continue training until an acceptable model is achieved.
*Deployment: We can run the trained model file directly from the built-in USB drive of the OpenMV.
We now demonstrate the specific process below:
01. Collect and upload the image dataset
1. First, log in to the website edgeimpulse.com, select “Log In”, enter the name of the new project, and select “Images” → “Classify multiple objects”

II. Select an image to upload. There are several ways to obtain the image dataset:
1. Capture images in real time using the OpenMV IDE. Use the “Dataset Editor” tool within the OpenMV IDE to create a new dataset and capture images in real time; for details, refer to our previous video tutorial on object classification and mask detection.

2. Images downloaded from the internet or captured using a mobile phone; preferably matching the actual environment in which detection will be performed. We have prepared approximately four to five hundred images of traffic signs, including “No Parking”, “No Horn”, and “Speed Limit 80 km/h”, all collected from real roads. This dataset is available for download from our GitHub repository and tutorial website.
02. Annotation Target
Upload approximately 100 images for each category; once the images have been uploaded, proceed to annotate them by marking both the location and category of the objects to be identified.


03. Train the Model
Configure the training parameters, change the resolution to 128×128, and select all default parameter settings; then click Save.

Generate features, with three colours representing three distinct markers.

Configure the target detection training parameters. The default number of training epochs is 60, the learning rate is 0.001, and the validation set ratio is 20%. Enable data augmentation. For transfer learning, you may select the default MobileNetV2 0.35 or MobileNetV2 0.1. (Note: SSD cannot be used, as it is not supported on OpenMV.)
Select Start Training. The neural network model obtained after training has an F1 Score of 91.2%, and we can save the current version.

Select Storage, enter your description, and select Save.

04. Deploy the model
Export the trained model file. Only the library needs to be exported; select OpenMV and click Build to begin deploying and exporting the model.

It will automatically download the model we exported, which in this case is the model trained on our 300 images, thereby completing the entire training process.

05. Run the Model
Running models on OpenMV:
Connect the OpenMV Plus, save the three trained files to the OpenMV’s built-in USB drive, and open the ei_object_detection.py file in the OpenMV IDE. Click Run, and the results will appear in the serial terminal.


This feature is also compatible with the OpenMV4. As the OpenMV4 does not have external SDRAM and has less memory, we can reduce the model size to enable it to run on the OpenMV4.
There are two methods to reduce the model’s size:
1. Reduce the training resolution from 128×128 to 96×96.
II. Modify the transfer learning model used, changing MobileNetV2 0.35 to MobileNetV2 0.1.
Finally, place the trained new file into the OpenMV4 and run it.
The above is our target point detection tutorial. We look forward to your results!

AI Sentinel Based on OpenMV: Automatic Alert for Unsecured Key Locations
Automatically detects whether the door of a key location has remained open for an extended period and issues a timely alert.

Personnel crossing boundary triggers an alarm; OpenMV interprets the “sense of security boundary”
Detects personnel or equipment crossing virtual perimeter lines.

“AI Urban Management” is here: an automated unauthorised street trading detection system built on OpenMV
Automatically identifies illegal street vending to support daily urban governance.

Misplaced gas cylinder? Use OpenMV to automatically trigger hazard alerts
Automatically identifies non-compliant placement of gas cylinders to proactively detect gas-related safety hazards.

Did you perform live-line work without wearing insulating gloves? OpenMV will immediately issue an alert!
Automatically identifies whether personnel are wearing insulating gloves to assist with safe operations.

Utilise the OpenMV smart camera to detect defects on aluminium plate surfaces in real time
Online identification of defects such as scratches and dents on aluminium plate surfaces to support quality control.