A real-time, multi-animal model for automatic face detection and identification of freely moving common marmosets based on YOLOv8 algorithms

  1. Jiayue Yang  Is a corresponding author
  2. James Wang
  3. Justine Cléry  Is a corresponding author
  1. The Neuro, Department of Neurology and Neurosurgery, McGill University, Canada
  2. Integrated Program in Neuroscience, McGill University, Canada
  3. McConnell Brain Imaging Centre, The Neuro, Montreal Neurological Institute and Hospital, McGill University, Canada
  4. Azrieli Centre for Autism Research, The Neuro, Canada
10 figures, 6 tables and 1 additional file

Figures

Figure 1 with 2 supplements
Training performance of the multi-marmoset face classification model for three adult marmosets, using the pre-trained YOLOv8 nano model.

(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 183. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP@50–95 (IoU = 0.5:0.95) for each label class.

Figure 1—figure supplement 1
Training performance of the multi-marmoset face classification model for three adult marmosets.

The validation distribution focal loss (DFL) of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 183.

Figure 1—video 1
Example annotations of the multi-marmoset face classification from three adult marmosets, from a test video excluded from model training and validation.

Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of adult marmoset faces and corresponding collar beads.

Normalized confusion matrix per-class classification across the six label classes in the adult marmoset recognition model.

The y-axis represents the predicted class, and the x-axis represents the manually labeled class. Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.

Figure 3 with 2 supplements
Training performance of the automatic facial and identity extraction model.

(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 124. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP@50–95 (IoU = 0.5:0.95) for each label class.

Figure 3—figure supplement 1
Training performance of the automatic facial and identity extraction model.

The validation distribution focal loss (DFL) of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 124.

Figure 3—video 1
Example annotations of the automatic face and identity extraction from two new marmosets at different ages.

Top panel: annotated output video of the extraction model. Bottom panel: timeline plot of the detected label classes of marmoset faces (blue) and collar (cyan) beads.

Normalized confusion matrix per-class classification across the two label classes in the automatic identity extraction model.

Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.

Figure 5 with 5 supplements
Training performance of the multi-marmoset face recognition model for young marmosets at 7 months.

(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 245. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP (IoU = 0.5:0.95) for each label class.

Figure 5—figure supplement 1
Training performance of the multi-marmoset face classification model for two young marmosets.

The validation distribution focal loss (DFL) of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 245.

Figure 5—video 1
Example annotations of the multi-marmoset face classification for Young1 marmoset at 7 months.

Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue), Young1 collar beads (cyan), Young2 (pink), and Young2 collar beads (lime green).

Figure 5—video 2
Example annotations of the multi-marmoset face classification for Young2 marmoset at 7 months.

Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue), Young1 collar beads (cyan), Young2 (pink), and Young2 collar beads (lime green).

Figure 5—video 3
Example annotations of the multi-marmoset face classification for Young1 marmoset at 11 months.

Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue), Young1 collar beads (cyan), Young2 (pink), and Young2 collar beads (lime green).

Figure 5—video 4
Example annotations of the multi-marmoset face classification for Young2 marmoset at 11 months.

Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue), Young1 collar beads (cyan), Young2 (pink), and Young2 collar beads (lime green).

Normalized confusion matrix per-class classification across the four label classes in the young marmoset recognition model.

Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.

Figure 7 with 5 supplements
Across-model visualization of the face similarity between marmoset pairs.

Four types of family relationships (mother-father, father-son, mother-son, and twin1-twin2) were compared, based on the training results of adult and young marmosets. The similarity was calculated using (A) cosine similarity and the (B) Euclidean distance. Heat map of the (C) cosine similarity and (D) Euclidean distance was plotted between the three relationship pairs in the adult family. The cosine similarity score ranged from 0 (very different) to 1 (exactly the same) for marmoset faces. The Euclidean distance score ranged from 0 (exactly the same) to 1 (very different) for marmoset faces.

Figure 7—figure supplement 1
Heat map of the cosine similarity between marmoset faces in the young family.

The two young marmosets (twins) were represented by their family, with the similarity score calculated between the two twins. The cosine similarity score ranged from 0 (very different) to 1 (exactly the same) for marmoset faces.

Figure 7—figure supplement 2
Heat map of the Euclidean distance between marmoset faces in the young family.

The two young marmosets (twins) were represented by their family, with the similarity score calculated between the two twins. The Euclidean distance score ranged from 0 (exactly the same) to 1 (very different) for marmoset faces.

Figure 7—video 1
Face-only annotations of the multi-marmoset face classification for Young1 marmoset at 16 months.

Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue) and Young2 (pink). The predictions of collar beads were excluded.

Figure 7—video 2
Face-only annotations of the multi-marmoset face classification for Young1 marmoset at 11 months.

Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue) and Young2 (pink). The predictions of collar beads were excluded.

Figure 7—video 3
Face-only annotations of the multi-marmoset face classification for Young2 marmoset at 11 months.

Top panel: annotated output video of the classifier. Bottom panel: timeline plot of the detected label classes of marmoset Young1 face (blue) and Young2 (pink). The predictions of collar beads were excluded.

Figure 8 with 1 supplement
Examples of marmoset face detection and classifiers for in-cage and open-field marmoset images.

This image illustrated the (A, B, D, E) detection of marmosets’ faces and (C) their identities across different cameras, environments, and marmoset behaviors. Additionally, it included scenarios where animals were housed in-cage (A–C) or moving freely (D–E). The open-field marmoset videos (D–E) were obtained from the YouTube Billabong Zoo, Koala, and Wildlife Park: https://www.youtube.com/watch?v=lDJNk4rjqqQ.

Figure 8—figure supplement 1
Identification performance of human experimenters across the five marmosets.

Identification accuracy of seven human experimenters across the five marmosets. The legend denoted the performance of the animal health technicians (blue) and the student/supervisor (pink).

Experimental setup of the camera system and primate chair for face detection and identification.

(A) Side view of the camera (green box) and primate chair, attached to the housing cage. (B) Frontal view of the camera fixed on the primate chair, enclosed within a protective cover. (C) A schematic illustration of the real-time facial image recording and automatic identification of the marmoset entering the primate chair.

Workflow and design of the marmoset facial detection and identification model.

(A) The architecture of the real-time marmoset facial recognition program. (B) Marmoset face images from three camera angles. (C) Bounding boxes of marmoset faces (green box) and collars (pink box) were manually labeled in the training and validation datasets to train the multi-marmoset face classification model. (D) The schematic of the automatic face (blue box) and collar bead (cyan box) extraction model.

Tables

Table 1
Performance comparison of multi-marmoset face classification models.

Comparison of recall, precision, F1 score, mean average precision (mAP) at intersection over union (IoU) = 0.5:0.95, validation distribution focal loss (DFL), and training time of the three pre-trained models, based on the performance of marmoset face detection and identification on the adult marmosets’ dataset. The highest values of each parameter were highlighted in bold.

MetricsYOLOv8 nanoYOLOv8 smallYOLOv8 medium
Recall0.9640.9570.956
Precision0.9320.9270.937
F1 score0.9480.9420.946
mAP@50–950.7100.7100.713
Val DFL0.9170.9531.010
Training time (s)5544590810705
Table 2
Statistical analysis of inter-individual face similarity (cosine similarity) between different family relationships, within the adult marmoset family.

Comparison of t-statistics, p-value, and Cohen’s d on cosine similarities were compared between the three family relationships. The relationship pairs that showed significant differences were bolded.

Relationshipst-Valuep-ValueCohen’s d
Mother-father vs. father-son–1.9410.083–0.868
Mother-father vs. mother-son–0.6130.548–0.274
Father-son vs. mother-son1.4600.1770.653
Table 3
Statistical analysis of inter-individual face similarity (Euclidean distance) between different family relationships, within the adult marmoset family.

Comparison of t-statistics, p-value, and Cohen’s d on Euclidean distances were compared between the three family relationships. The relationship pairs that showed significant differences were bolded.

Relationshipst-Valuep-ValueCohen’s d
Mother-father vs. father-son2.2760.0461.018
Mother-father vs. mother-son0.7470.4650.334
Father-son vs. mother-son–1.3970.192–0.625
Appendix 1—table 1
Model performance evaluation of the validation dataset in the adult marmoset recognition model.

Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.

Label classesPrecisionRecallmAP50mAP@50–95
All0.9360.9550.9800.713
Adult10.8290.9450.9570.841
Adult20.9260.9830.9910.787
Adult30.9510.9980.9920.868
collar_Adult10.9680.9660.9880.603
collar_Adult20.9780.9720.9900.665
collar_Adult30.9620.8660.9590.515
Appendix 1—table 2
Model performance evaluation of the validation dataset in the automatic face and identity extraction model.

Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.

Label classesPrecisionRecallmAP50mAP@50–95
All0.9380.9680.9840.715
Face0.9190.9860.9900.831
Collar0.9570.9500.9780.600
Appendix 1—table 3
Model performance evaluation of the validation dataset in the young marmoset recognition model.

Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.

Label classesPrecisionRecallmAP50mAP@50–95
All0.9790.9750.9940.890
Young11.0000.9350.9930.950
collar_Young10.9981.0000.9950.808
Young20.9581.0000.9950.947
collar_Young20.9630.9640.9930.855

Additional files

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Jiayue Yang
  2. James Wang
  3. Justine Cléry
(2026)
A real-time, multi-animal model for automatic face detection and identification of freely moving common marmosets based on YOLOv8 algorithms
eLife 15:RP110932.
https://doi.org/10.7554/eLife.110932.3