A real-time, multi-animal model for automatic face detection and identification of freely moving common marmosets based on YOLOv8 algorithms
Figures
Training performance of the multi-marmoset face classification model for three adult marmosets, using the pre-trained YOLOv8 nano model.
(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 183. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP@50–95 (IoU = 0.5:0.95) for each label class.
Training performance of the multi-marmoset face classification model for three adult marmosets.
The validation distribution focal loss (DFL) of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 183.
Example annotations of the multi-marmoset face classification from three adult marmosets, from a test video excluded from model training and validation.
Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of adult marmoset faces and corresponding collar beads.
Normalized confusion matrix per-class classification across the six label classes in the adult marmoset recognition model.
The y-axis represents the predicted class, and the x-axis represents the manually labeled class. Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.
Training performance of the automatic facial and identity extraction model.
(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 124. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP@50–95 (IoU = 0.5:0.95) for each label class.
Training performance of the automatic facial and identity extraction model.
The validation distribution focal loss (DFL) of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 124.
Example annotations of the automatic face and identity extraction from two new marmosets at different ages.
Top panel: annotated output video of the extraction model. Bottom panel: timeline plot of the detected label classes of marmoset faces (blue) and collar (cyan) beads.
Normalized confusion matrix per-class classification across the two label classes in the automatic identity extraction model.
Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.
Training performance of the multi-marmoset face recognition model for young marmosets at 7 months.
(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 245. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP (IoU = 0.5:0.95) for each label class.
Training performance of the multi-marmoset face classification model for two young marmosets.
The validation distribution focal loss (DFL) of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 245.
Example annotations of the multi-marmoset face classification for Young1 marmoset at 7 months.
Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue), Young1 collar beads (cyan), Young2 (pink), and Young2 collar beads (lime green).
Example annotations of the multi-marmoset face classification for Young2 marmoset at 7 months.
Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue), Young1 collar beads (cyan), Young2 (pink), and Young2 collar beads (lime green).
Example annotations of the multi-marmoset face classification for Young1 marmoset at 11 months.
Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue), Young1 collar beads (cyan), Young2 (pink), and Young2 collar beads (lime green).
Example annotations of the multi-marmoset face classification for Young2 marmoset at 11 months.
Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue), Young1 collar beads (cyan), Young2 (pink), and Young2 collar beads (lime green).
Normalized confusion matrix per-class classification across the four label classes in the young marmoset recognition model.
Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.
Across-model visualization of the face similarity between marmoset pairs.
Four types of family relationships (mother-father, father-son, mother-son, and twin1-twin2) were compared, based on the training results of adult and young marmosets. The similarity was calculated using (A) cosine similarity and the (B) Euclidean distance. Heat map of the (C) cosine similarity and (D) Euclidean distance was plotted between the three relationship pairs in the adult family. The cosine similarity score ranged from 0 (very different) to 1 (exactly the same) for marmoset faces. The Euclidean distance score ranged from 0 (exactly the same) to 1 (very different) for marmoset faces.
Heat map of the cosine similarity between marmoset faces in the young family.
The two young marmosets (twins) were represented by their family, with the similarity score calculated between the two twins. The cosine similarity score ranged from 0 (very different) to 1 (exactly the same) for marmoset faces.
Heat map of the Euclidean distance between marmoset faces in the young family.
The two young marmosets (twins) were represented by their family, with the similarity score calculated between the two twins. The Euclidean distance score ranged from 0 (exactly the same) to 1 (very different) for marmoset faces.
Face-only annotations of the multi-marmoset face classification for Young1 marmoset at 16 months.
Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue) and Young2 (pink). The predictions of collar beads were excluded.
Face-only annotations of the multi-marmoset face classification for Young1 marmoset at 11 months.
Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue) and Young2 (pink). The predictions of collar beads were excluded.
Face-only annotations of the multi-marmoset face classification for Young2 marmoset at 11 months.
Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue) and Young2 (pink). The predictions of collar beads were excluded.
Examples of marmoset face detection and classifiers for in-cage and open-field marmoset images.
This image illustrated the (A, B, D, E) detection of marmosets’ faces and (C) their identities across different cameras, environments, and marmoset behaviors. Additionally, it included scenarios where animals were housed in-cage (A–C) or moving freely (D–E). The open-field marmoset videos (D–E) were obtained from the YouTube Billabong Zoo, Koala, and Wildlife Park: https://www.youtube.com/watch?v=lDJNk4rjqqQ.
Identification performance of human experimenters across the five marmosets.
Identification accuracy of seven human experimenters across the five marmosets. The legend denoted the performance of the animal health technicians (blue) and the student/supervisor (pink).
Experimental setup of the camera system and primate chair for face detection and identification.
(A) Side view of the camera (green box) and primate chair, attached to the housing cage. (B) Frontal view of the camera fixed on the primate chair, enclosed within a protective cover. (C) A schematic illustration of the real-time facial image recording and automatic identification of the marmoset entering the primate chair.
Workflow and design of the marmoset facial detection and identification model.
(A) The architecture of the real-time marmoset facial recognition program. (B) Marmoset face images from three camera angles. (C) Bounding boxes of marmoset faces (green box) and collars (pink box) were manually labeled in the training and validation datasets to train the multi-marmoset face classification model. (D) The schematic of the automatic face (blue box) and collar bead (cyan box) extraction model.
Tables
Performance comparison of multi-marmoset face classification models.
Comparison of recall, precision, F1 score, mean average precision (mAP) at intersection over union (IoU) = 0.5:0.95, validation distribution focal loss (DFL), and training time of the three pre-trained models, based on the performance of marmoset face detection and identification on the adult marmosets’ dataset. The highest values of each parameter were highlighted in bold.
| Metrics | YOLOv8 nano | YOLOv8 small | YOLOv8 medium |
|---|---|---|---|
| Recall | 0.964 | 0.957 | 0.956 |
| Precision | 0.932 | 0.927 | 0.937 |
| F1 score | 0.948 | 0.942 | 0.946 |
| mAP@50–95 | 0.710 | 0.710 | 0.713 |
| Val DFL | 0.917 | 0.953 | 1.010 |
| Training time (s) | 5544 | 5908 | 10705 |
Statistical analysis of inter-individual face similarity (cosine similarity) between different family relationships, within the adult marmoset family.
Comparison of t-statistics, p-value, and Cohen’s d on cosine similarities were compared between the three family relationships. The relationship pairs that showed significant differences were bolded.
| Relationships | t-Value | p-Value | Cohen’s d |
|---|---|---|---|
| Mother-father vs. father-son | –1.941 | 0.083 | –0.868 |
| Mother-father vs. mother-son | –0.613 | 0.548 | –0.274 |
| Father-son vs. mother-son | 1.460 | 0.177 | 0.653 |
Statistical analysis of inter-individual face similarity (Euclidean distance) between different family relationships, within the adult marmoset family.
Comparison of t-statistics, p-value, and Cohen’s d on Euclidean distances were compared between the three family relationships. The relationship pairs that showed significant differences were bolded.
| Relationships | t-Value | p-Value | Cohen’s d |
|---|---|---|---|
| Mother-father vs. father-son | 2.276 | 0.046 | 1.018 |
| Mother-father vs. mother-son | 0.747 | 0.465 | 0.334 |
| Father-son vs. mother-son | –1.397 | 0.192 | –0.625 |
Model performance evaluation of the validation dataset in the adult marmoset recognition model.
Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.
| Label classes | Precision | Recall | mAP50 | mAP@50–95 |
|---|---|---|---|---|
| All | 0.936 | 0.955 | 0.980 | 0.713 |
| Adult1 | 0.829 | 0.945 | 0.957 | 0.841 |
| Adult2 | 0.926 | 0.983 | 0.991 | 0.787 |
| Adult3 | 0.951 | 0.998 | 0.992 | 0.868 |
| collar_Adult1 | 0.968 | 0.966 | 0.988 | 0.603 |
| collar_Adult2 | 0.978 | 0.972 | 0.990 | 0.665 |
| collar_Adult3 | 0.962 | 0.866 | 0.959 | 0.515 |
Model performance evaluation of the validation dataset in the automatic face and identity extraction model.
Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.
| Label classes | Precision | Recall | mAP50 | mAP@50–95 |
|---|---|---|---|---|
| All | 0.938 | 0.968 | 0.984 | 0.715 |
| Face | 0.919 | 0.986 | 0.990 | 0.831 |
| Collar | 0.957 | 0.950 | 0.978 | 0.600 |
Model performance evaluation of the validation dataset in the young marmoset recognition model.
Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.
| Label classes | Precision | Recall | mAP50 | mAP@50–95 |
|---|---|---|---|---|
| All | 0.979 | 0.975 | 0.994 | 0.890 |
| Young1 | 1.000 | 0.935 | 0.993 | 0.950 |
| collar_Young1 | 0.998 | 1.000 | 0.995 | 0.808 |
| Young2 | 0.958 | 1.000 | 0.995 | 0.947 |
| collar_Young2 | 0.963 | 0.964 | 0.993 | 0.855 |