A real-time, multi-animal model for automatic face detection and identification of freely moving common marmosets based on YOLOv8 algorithms

  1. Jiayue Yang  Is a corresponding author
  2. James Wang
  3. Justine Cléry  Is a corresponding author
  1. The Neuro, Department of Neurology and Neurosurgery, McGill University, Canada
  2. Integrated Program in Neuroscience, McGill University, Canada
  3. McConnell Brain Imaging Centre, The Neuro, Montreal Neurological Institute and Hospital, McGill University, Canada
  4. Azrieli Centre for Autism Research, The Neuro, Canada

eLife Assessment

This valuable study presents a real-time system for identifying multiple unrestrained marmosets in a home cage setting using a combination of facial features and color-coded beads. While there is solid evidence that the system has a precision comparable to human experimenters in the tested scenarios, there is limited evidence that this would generalize to unconstrained multi-animal environments

https://doi.org/10.7554/eLife.110932.3.sa0

Abstract

Precise and up-to-date information about animal location and identity allows us to better quantify individual behaviors in studies of neural activity, cognition, and animal health. In socially housed laboratory animals, identification is usually defined by observation or invasive markers, making the data collection time-consuming, variable across experimenters, and disruptive to animals. We established an automatic pipeline for real-time identification of common marmosets in captivity using a close-view camera. It uses the supervised deep-learning YOLOv8 model to localize individuals, detect faces, and classify identities. Moreover, we use recognition of uniquely color-coded collar beads to improve detection accuracy among visually similar individuals. Across adult and juvenile marmosets, our system automatically identifies marmosets with >82.9% precision and >91.5% recall, achieving human-level performance. This pipeline is designed to be easy to use and generalizable across non-human primate species, ages, and recording hardware, providing rapid and automatic identity recognition from real-time video.

Introduction

In behavioral neuroscience, an accurate identification of laboratory animals is crucial for welfare assessment and ethical care (National Research Council, 2011; Vidal et al., 2021). More importantly, the animal identity is also required when conducting individual-specific behavioral experiments. However, identifying specific animals in a family group can be challenging, especially if they are not marked and share similar facial features. Invasive recognition methods, including branding, tattooing, and ear tagging, may disrupt animals’ behaviors and cause potential stress (Carstens and Moberg, 2000; Lim et al., 2019; Roughan and Sevenoaks, 2019). To avoid such harms, video analysis is becoming more popular as a non-invasive method to identify the animals by their appearance, with no direct contact or handling being involved (Norouzzadeh et al., 2018; Schindler and Steinhage, 2021). However, collecting video data of animals that is subsequently labeled by human observation can be time-consuming and hard to reproduce across experimenters (Buchan et al., 2003; Marion et al., 2020). Therefore, considering replacement of human work, machine learning tools have been introduced and increasingly used, providing an accurate and automatic estimation of animal identities.

Computer vision and machine learning have made many advancements and offered the foundation to build these automatic systems (Sturman et al., 2020). Within these approaches, deep learning, a subset of machine learning, has been widely used in animal identification based on video and image data. Specifically, by utilizing deep learning with video recordings, object detection allows an accurate localization and identification of animals, even with those who live in complicated environments (Yu et al., 2018; Zhuang et al., 2025). Algorithms such as ResNet (He et al., 2015) and You Only Look Once (YOLO) models Jocher et al., 2023; Redmon et al., 2015 have demonstrated promising accuracy and computational speed in identifying animals in naturalistic environments (Bakana et al., 2024; Petso et al., 2021). In addition, unique facial features can also be used as inputs to classify and recognize animals. This approach has the advantage that it provides information regarding whether a specific animal is present or absent in the camera recording view at each time point. Although facial recognition has been demonstrated to be successful in wild and domestic animals (Bergman et al., 2024; Norouzzadeh et al., 2021; Schofield et al., 2023), its application in controlled laboratory settings is not yet extensively evaluated.

The common marmoset (Callithrix jacchus), a small non-human primate, is becoming more popular recently in neuroscience research for many reasons (Kishi et al., 2014; Okano, 2021; Sasaki et al., 2009). Marmosets are family-bonded, cooperative, and notably pro-social. Moreover, they can perform multiple series of behavioral and cognitive tasks, such as observational and reversal learning (Koski and Burkart, 2015; Miller et al., 2016). As marmosets are typically maintained in social family groups (Yoshimoto et al., 2018), accurate identification of specific individuals is therefore essential for assessing which animal is performing a given task, which also enables tracking of their behavioral performance over time. However, existing identification techniques of marmosets usually involve wearable micro-sensing devices, requiring stable device positioning to the sensor and constant adjustments to ensure the devices fit the animal. This automatic system can also become unreliable during rapid movement or if multiple marmosets are present at the sensor, which limits the consistency and stability of identity tracking in freely moving marmosets during their tasks. Thus, facial recognition offers a non-invasive alternative, without a strict need for a physical identity marker, enabling continuous, accurate, and automated identification when the marmoset enters the designated task performance space.

Here, we developed a pipeline for automatically detecting and identifying marmosets simultaneously from real-time videos, based on their faces. Previous studies have introduced the feasibility of the methods we used in this study (Dave et al., 2023; XiaoAn et al., 2024), of which we adapted them for facial-based marmoset recognition. We apply the YOLOv8 model (Jocher et al., 2023), trained on images of specific individuals (three adult and two young marmosets), to generate real-time marmoset recognition on a frame-by-frame manner. In addition to marmoset faces, we use the color-coded bead on individual collars to facilitate the identity prediction accuracy. Finally, we further evaluated our model performance on a dataset from both adult marmosets and young marmosets at different developmental stages. We show that our facial-based pipeline achieves high detection performance comparable to human expert level, with strong generalizability across recording settings and marmosets at various ages.

Results

Comparison of pre-trained YOLOv8 models

Model performance included the precision (the proportion of correct positive predictions), recall (the proportion of corrected predicted ground truth labels), and mAP@50–95 (the average detection accuracy across different intersection over union [IoU] object localization thresholds; see Materials and methods section for detailed definitions). The results of the training and evaluation of the pre-trained YOLOv8 models (YOLOv8 nano, YOLOv8 small, YOLOv8 medium) were shown in Table 1. The models reached their best performance at 180 epochs (YOLOv8 nano), 132 epochs (YOLOv8 small), and 124 epochs (YOLOv8 medium). Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs). By comparing the performance differences between the YOLOv8 pre-trained networks, we aimed to find the most suitable model for the real-time marmoset identification program.

Table 1
Performance comparison of multi-marmoset face classification models.

Comparison of recall, precision, F1 score, mean average precision (mAP) at intersection over union (IoU) = 0.5:0.95, validation distribution focal loss (DFL), and training time of the three pre-trained models, based on the performance of marmoset face detection and identification on the adult marmosets’ dataset. The highest values of each parameter were highlighted in bold.

MetricsYOLOv8 nanoYOLOv8 smallYOLOv8 medium
Recall0.9640.9570.956
Precision0.9320.9270.937
F1 score0.9480.9420.946
mAP@50–950.7100.7100.713
Val DFL0.9170.9531.010
Training time (s)5544590810705

As presented in Table 1, the YOLOv8 nano model reached the best overall detection performance of the adult marmosets’ dataset. The YOLOv8 medium model achieved a higher precision among all three pre-trained networks. However, the recall and F1 score were higher for the YOLOv8 nano model, indicating robust model sensitivity and a balanced precision-recall value. Mean average precision (mAP@50–95) showed values ranging from 0.710 to 0.713 for the three models. The YOLOv8 medium model reached the highest detection accuracy for adult marmoset faces, though the difference was minor while comparing to the YOLOv8 nano and small models. In addition, we considered the validation distribution focal loss (DFL) and training time to support the selection of the optimal pre-trained model. Validation DFL increased with model sizes, suggesting that the detection localization became less accurate from YOLOv8 nano to YOLOv8 medium. Moreover, the training time increased from YOLOv8 nano to the medium model, while the YOLOv8 medium took about twice the training time compared to the YOLOv8 nano model. Thus, we selected the YOLOv8 nano model for training the real-time program due to its high computational speed and efficiency, while maintaining reliable marmoset face prediction compared to larger YOLOv8 models (Jocher et al., 2023).

Prediction based on adult marmoset faces shows reliable identification

We applied the YOLOv8 nano model for the adult marmoset face classification. The training curves of the overall precision, recall, and mAP@50–95 (intersection over union [IoU]=0.5:0.95) were visualized in Figure 1A, B, C, Figure 1—figure supplement 1. The face classifier weights were selected at the training epoch 183 (denoted by the red dotted line), reaching the precision at 0.932, recall at 0.964, and the mAP@50–95 (IoU = 0.5:0.95) at 0.710 (see Table 1). We evaluated the face classifier performance per label class using the adult marmoset’ validation dataset (Figure 1D, see Appendix 1—table 1). Analyzing the label classes, the adult marmoset face classifier showed stable and high precision and recall across both individual identity (i.e. Adult1, Adult2, Adult3) and bead-color (i.e. collar_Adult1, collar_Adult2, collar_Adult3) label classes (Figure 1—video 1). In addition, the marmoset face classes showed high mAP across varied IoU thresholds, suggesting a robust prediction localization for marmoset faces (mAP@50–95=0.841, 0.787, 0.868 for three adults, respectively). The lower mAP value observed for the bead-color label category (mAP@50–95=0.603, 0.665, 0.515 for collars of three adults) suggested a greater variability in bounding box localization for the marmoset collar beads. This effect was consistent with the small bounding box sizes of the collar-bead label class, which increased the model sensitivity to localization variability in the prediction (example detection shown in Figure 10C). The normalized confusion matrix comparison was shown in Figure 2. The detection accuracy shows similar performance using the adult marmosets’ training dataset (Figure 2A) and validation dataset (Figure 2B), with very low occurrence of false detection between the marmoset individuals and the background images.

Figure 1 with 2 supplements see all
Training performance of the multi-marmoset face classification model for three adult marmosets, using the pre-trained YOLOv8 nano model.

(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 183. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP@50–95 (IoU = 0.5:0.95) for each label class.

Normalized confusion matrix per-class classification across the six label classes in the adult marmoset recognition model.

The y-axis represents the predicted class, and the x-axis represents the manually labeled class. Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.

Automatic facial and identity extraction accurately localizes new marmoset faces and their collar beads

The labeled adult marmosets’ dataset was used for training the automatic facial and identity extraction, using the YOLOv8 nano pre-trained model. This framework aimed to extract and localize the marmoset faces and collar beads from the collected video, including new marmosets. The trend of the training was shown in Figure 3A–C and Figure 3—figure supplement 1, where the detection accuracy improved rapidly in the beginning, with the increasing rate then stabilizing and reaching a plateau. We selected the final weights at the training epoch of 124, with a prediction score at 0.940, a recall score at 0.970, and the mAP@50–95 (IoU = 0.5:0.95) at 0.716. In addition, precision, recall, and mAP@50–95 (IoU = 0.5:0.95) were calculated per label class for marmoset faces and collar beads (Figure 3D). Based on the model performance of the adult marmosets’ validation dataset, the final face extraction system accurately predicted the marmoset faces (precision = 0.919, recall = 0.968) and collar beads (precision = 0.957, recall = 0.95). The bead-color class showed lower mAP@50–95 (IoU = 0.5:0.95) values at 0.6, compared to the mAP@50–95 (IoU = 0.5:0.95) for marmoset face class at 0.95, suggesting a similar effect for the face classification of adult marmosets in Figure 1D and Appendix 1—table 2. The normalized confusion matrices were compared between the training and validation dataset (Figure 4). It was evident that the automatic face and identity extraction exhibited a high detection accuracy. The background class of the confusion matrix showed frequent predictions as marmoset faces and collar beads for the training (Figure 4A) and validation set (Figure 4B). However, it does not necessarily indicate incorrect predictions or misclassifications. Instead, these values were mostly explained by multiple detections of the same object class. For instance, additional marmoset faces were predicted when multiple animals were present within a single video frame. The long collar structure or motion blur of the marmosets could also cause multiple detections of beads that belong to the same collar. This also corresponded to the high precision and recall scores observed across prediction classes (Figure 3D), suggesting that the increased background false positives were mainly related to the object-count discrepancies, instead of poor detection performance.

Figure 3 with 2 supplements see all
Training performance of the automatic facial and identity extraction model.

(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 124. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP@50–95 (IoU = 0.5:0.95) for each label class.

Normalized confusion matrix per-class classification across the two label classes in the automatic identity extraction model.

Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.

To test the performance of the automatic face and identity extraction, we used different short clips of new marmosets. We showed that our automatic face and identity extraction pipeline could detect and localize new marmoset faces and collar beads at different ages. Figure 3—video 1 showed an example of the detection of faces and bead-color labels for two new younger marmosets at 7 months and 17 months, with the confidence threshold of the detection set at 0.6.

Comparison of identification accuracy in young marmosets across developmental stages

The training and validation datasets of the young marmosets were annotated automatically using the automatic facial and identity extraction. All automatically generated labels were reviewed by the experimenter, and mislabeled images were removed prior to the model training. The training performance curves were presented in Figure 5A–C and Figure 5—figure supplement 1. The final weights were selected at the training epoch of 245, where the training performance metrics achieved a plateau. This program achieved the best performance metrics at this epoch, with precision score at 0.979, recall at 0.975, and mAP@50–95 (IoU = 0.5:0.95) at 0.892. In the young marmoset faces and their collar bead label classes, our face classifier reached high precision and recall performances using the validation dataset (Figure 5D, see Appendix 1—table 3). The precision and recall scores for the label classes were: (1) face of Young1: precision = 1, recall = 0.935; (2) face of Young2: precision = 0.958, recall = 1; (3) collar of Young1: precision = 0.998, recall = 1; (4) collar of Young2: precision = 0.963, recall = 0.964. In terms of localization accuracy, we found a similar pattern of reduced mAP@50–95 (IoU = 0.5:0.95) in bead-color class (mAP@50–95=0.808, 0.855), compared to the marmoset faces (mAP@50–95=0.95, 0.947). We further computed the normalized confusion matrices of this face classifier using the training and validation dataset. We found that our face classification was relatively accurate across detection of most class labels (Figure 6). The normalized confusion matrices showed high accuracy and consistency of most marmoset faces and collars detection in training (Figure 6A) and validation (Figure 6B) tests, with some exceptions. Particularly, the background was frequently identified as the collar of Young2 marmoset. This elevated background score was likely contributed by the multi-color design of the Young2 marmoset collar, making it more difficult to distinguish compared to collars with a single bead color. In this occasion, if one bead is occluded, blurred, or outside the field of view, the other visible collar bead could affect the prediction and lead to an incorrect identification from the ground truth.

Figure 5 with 5 supplements see all
Training performance of the multi-marmoset face recognition model for young marmosets at 7 months.

(A) Precision of all detection classes across training epochs. The red dotted line denoted the final model with the best performance at training epoch 245. (B) Similar to (A), except for overall recall. (C) Similar to (A), except for the mean average precision (mAP) at the intersection over union (IoU) at 0.5:0.95. (D) Overall model precision, recall, and the mAP (IoU = 0.5:0.95) for each label class.

Normalized confusion matrix per-class classification across the four label classes in the young marmoset recognition model.

Proportion was generated by the (A) training dataset and the (B) validation dataset, showing whether certain classes were frequently mislabeled as a different class.

To test the detection performance of the face classifier trained on 7-month-old young marmosets across different developmental stages, we used short clips of the two young marmosets at the age of 7 months (Figure 5—video 1, Figure 5—video 2) and 11 months (Figure 5—video 3, Figure 5—video 4). Figure 5—videos 1–4 presented the example detections of the young twins at different ages, showing the difficulty of identifying and distinguishing between twin marmosets.

By combining the identification of marmoset faces and collar beads, the program showed valid detection of the correct identity of individual marmosets. However, the detection of the twin marmosets also posed various challenges for computer vision and resulted in mislabeling, including highly similar faces of twin marmosets, face blur or mislabeling due to complex backgrounds and dark lighting conditions (Figure 5—video 1), fur occluding collar beads, marmosets only showing faces during short time intervals (Figure 5—video 4), similar or same color of collar beads between two marmosets (Figure 5—video 3, Figure 5—video 4), and fewer training data (young marmosets’ model only included 449 training and validation images for two marmosets at 7 months of age).

We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Figure 7—video 1, Figure 7—video 2, Figure 7—video 3). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Figure 7—video 1).

Facial similarities of marmosets affect the performance of the facial detection and identification model

We quantified the inter-individual face similarities using cosine similarity and Euclidean distance between mean normalized embeddings of marmoset pairs. We included both adult and young marmosets’ dataset and computed the face similarity based on the four relationship pairs: (1) mother-father, (2) father-son, (3) mother-son, and (4) twin-twin. To visualize across the two independently trained models, the inter-individual face similarity scores were normalized using z-score within each model (Figure 7, Figure 7—figure supplements 1 and 2).

Figure 7 with 5 supplements see all
Across-model visualization of the face similarity between marmoset pairs.

Four types of family relationships (mother-father, father-son, mother-son, and twin1-twin2) were compared, based on the training results of adult and young marmosets. The similarity was calculated using (A) cosine similarity and the (B) Euclidean distance. Heat map of the (C) cosine similarity and (D) Euclidean distance was plotted between the three relationship pairs in the adult family. The cosine similarity score ranged from 0 (very different) to 1 (exactly the same) for marmoset faces. The Euclidean distance score ranged from 0 (exactly the same) to 1 (very different) for marmoset faces.

Within the model trained using the adult marmoset family, both face similarity measures showed a consistent pattern across the family relationships. In the adult family, we found that the father-son pair showed the highest face similarity values, with the cosine similarity at a positive mean z-score (z=0.43) and the Euclidean distance at a negative mean z-score (z=–0.47) (Figure 7C and D). This pattern suggested that the father-son pair exhibited a higher-than-average similarity relative to other family relationships. On the other hand, the mother-father relationship showed the lowest face similarity in the family (negative cosine similarity z=–0.38, positive Euclidean distance z=0.43) compared to the adult marmosets’ model mean. The similarity of the mother-son pair was calculated at cosine similarity z=–0.05 and Euclidean distance z=0.04, indicating an intermediate face similarity in the adult family.

The young marmosets’ identifier was independently trained specifically on the two twin marmosets. Thus, the normalized face similarity z-scores were centered near zero and not informative of the twin similarity in family relationship comparison (Figure 7). However, we found a narrow distribution of the image embeddings between the twin marmosets, indicating nearly invariant facial embeddings and structure between the two twins, with an extremely low raw variance for cosine similarity (std = 0.0005) and Euclidean distance (std = 0.0034).

We performed the statistical tests only on the face classifier for the adult marmoset family, as the twin marmoset model only involved two individuals and thus was not valid for within-model statistical analysis (Tables 2 and 3). Cosine similarity showed a trend of lower similarity in the mother-father pair compared to the father-son pair (t=–1.941, p=0.083>0.05, Cohen’s d=–0.868 [large effect]), though not significant. Differences between other relationship pairs were not significant (Table 2). Our analysis on the Euclidean distance indicated a significant difference between mother-father and father-son pairs (t=2.28, p=0.046<0.05, d=1.02 [large effect]), while the remaining relationship pairs appeared non-significant (Table 3).

Table 2
Statistical analysis of inter-individual face similarity (cosine similarity) between different family relationships, within the adult marmoset family.

Comparison of t-statistics, p-value, and Cohen’s d on cosine similarities were compared between the three family relationships. The relationship pairs that showed significant differences were bolded.

Relationshipst-Valuep-ValueCohen’s d
Mother-father vs. father-son–1.9410.083–0.868
Mother-father vs. mother-son–0.6130.548–0.274
Father-son vs. mother-son1.4600.1770.653
Table 3
Statistical analysis of inter-individual face similarity (Euclidean distance) between different family relationships, within the adult marmoset family.

Comparison of t-statistics, p-value, and Cohen’s d on Euclidean distances were compared between the three family relationships. The relationship pairs that showed significant differences were bolded.

Relationshipst-Valuep-ValueCohen’s d
Mother-father vs. father-son2.2760.0461.018
Mother-father vs. mother-son0.7470.4650.334
Father-son vs. mother-son–1.3970.192–0.625

A real-time marmoset identification program based on the trained networks

We developed a real-time interface for detecting and identifying the marmosets, based on the trained models of multi-marmoset classification. The real-time footages were acquired using the same camera and experimental setup as the training videos (see examples in Figure 1—video 1–Figure 5—video 4). Our real-time program processes the real-time footage like the detection of the offline collected videos, using the best model selected from each multi-marmoset face classification training model. Prior to running the real-time program, the experimenter first trains the multi-marmoset face classification model based on the specific subjects of interest. Next, the model with the best performance from training is selected and inputted into the real-time program. No programming is required from the experimenter using this real-time marmoset face identification program. The experimenter clicks a run button on the program console to initiate the real-time detection. While this program is running, the experimenter can check the frame-by-frame label detection results (30 frames per second) displayed on the screen. Potential mislabeling is mitigated by combining the marmoset faces and collar beads detection, with greater weights assigned to the collar bead detections than the marmoset face (example mislabeling in Figure 5—videos 3 and 4). The program automatically records the most frequently detected marmoset identity across 30 frames (approximately 1 s). At the end of the experiment, the program is terminated by clicking a stop button of the Python console. After the completion of the experiment, the full record of marmoset identity detection results is available for review by the experimenter.

Discussion

We developed a real-time computer vision program for automatic marmoset identity recognition, using their facial features and uniquely color-coded bead collars. By combining the automatic annotation of the real-time footage with the individual-specific classification (i.e. faces and collar beads), our program allows continuous identity tracking during behavioral experiments, with minimal human interference. Utilizing the object detection YOLOv8 model based on pre-trained image networks, we adapted it to a non-human primate (i.e. marmosets) face dataset and applied it in real-time tracking. Pose estimation tools have been widely used in characterizing animal behavior; however, these tools require extensive computational power, thus limiting the identification between visually similar individuals, especially animals housed in family units (Camilleri et al., 2023; Gill et al., 2025; Lauer et al., 2022; Mathis et al., 2018; Nath et al., 2019). Notably, the main goal of the current system is marmoset identity recognition, instead of pose estimation or behavioral analysis, such that identity characterization is not dependent on posture cues or movement tracking. Thus, we selected the YOLOv8 object detection algorithms for the development of our real-time marmoset face recognition pipeline. Among the YOLOv8 pre-trained models, we selected the lightweight YOLOv8 nano model, which provided the optimal balance between detection accuracy and computational inference speed, supporting its feasibility for the final real-time marmoset identity recognition program.

Our real-time marmoset recognition pipeline includes two face classifications for both adult and young marmosets, with an additional automatic face and identity extraction program for all marmosets. The models presented in this paper achieved reliable detection accuracy across adult marmosets’ and twin marmosets’ datasets, with detection accuracy improving with increased amounts of training images. While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition. Moreover, we anticipate that experimenters can integrate this program into behavioral experiments, as an automated marmoset identity extraction and classification tool for real-time video monitoring. Once fully trained on recognizing subject marmosets, this pipeline operates automatically and can be applied across individuals of different ages, with minimal manual work needed. While existing marmoset identification approaches usually utilize visible markers, radio frequency identification (RFID), or observation, the manual work and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters check the identity continuously, instead of outperforming the experimenters’ role. Its automated pipeline substantially reduces the time and work required for traditional manual identity labeling, while maintaining an expert-level human performance and reproducibility across experimenters (95.83% average accuracy for animal health technicians, responsible for daily health checks and husbandry, while lab experiments range between 25% and 80% of accuracy depending on the amount of time spent with each animal, Figure 8—figure supplement 1). The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.

The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger fields of view or phone cameras (Figure 8). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair systems, and computer-based devices. Based on the collected individual marmoset face data, the automatic extraction program allows consistent identity annotation of marmoset faces and collar beads, ensuring accurate and stable identity interpretation during the marmoset face classification model training. The experimenter only needs to run the Python pipeline scripts to perform the real-time marmoset identification during experiments, without the need to manually label large marmoset identity datasets. Moreover, the system operates directly based on the raw video input from the camera, with no preprocessing such as video cropping or resolution modification required. The identity recognition system recognizes each marmoset based on pre-defined labels from collected videos of individuals, without other external tools such as RFID systems (Pereira et al., 2023). We focused the marmoset identification on individual-specific facial features and uniquely colored collar beads to ensure the pipeline’s robustness across various recording conditions, regardless of lighting conditions, camera angles, or animal postures. Together, this pipeline design and setup enable fast, reliable, and accurate identity recognition for efficient real-time monitoring of multiple animals in complex experimental environments. However, with the high visual similarity between closely related marmoset family members, the facial features and the collar beads must be carefully integrated into our pipeline design, as even experienced human observers can misidentify marmoset twins.

Figure 8 with 1 supplement see all
Examples of marmoset face detection and classifiers for in-cage and open-field marmoset images.

This image illustrated the (A, B, D, E) detection of marmosets’ faces and (C) their identities across different cameras, environments, and marmoset behaviors. Additionally, it included scenarios where animals were housed in-cage (A–C) or moving freely (D–E). The open-field marmoset videos (D–E) were obtained from the YouTube Billabong Zoo, Koala, and Wildlife Park: https://www.youtube.com/watch?v=lDJNk4rjqqQ.

Because of the high extent of natural facial similarity between related marmosets, our real-time pipeline can experience challenges in identifying closely related individuals, as the results in Figures 5—7 and Figure 5—videos 1–4 indicate. Fur textures, coloring patterns, and genetic relatedness contribute to the morphological similarities, making it challenging for human observers when identifying individual primates (Alvergne et al., 2009; Guan et al., 2023; Leopold and Rhodes, 2010). Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Figure 5—videos 3 and 4) and beads that are occluded by fur (Figure 5—videos 1–4). In addition, collar beads may change over time or not be worn by all animals. An alternate experimental solution is to improve collar visibility, including using distinct color codes across individuals within a family, increasing the size of the beads, or increasing the number of collar beads to reduce occlusion. Moreover, the performance of the system is strongly dependent on the amount and variability of the training data, with identity classification improving as more marmoset images are involved in the model training. This relation is specifically important when characterizing between young and adult marmosets, as facial features may change across marmosets’ developmental stages. Images from both the juvenile and adult periods of the same marmoset can be included in the model training, which could improve model performance and generalizability across developmental stages, and avoid repeated training of the individual recognition model at different ages.

Because our pipeline was designed specifically for marmoset identity recognition, it was optimized for characterizing individuals within a defined housing unit, which usually represents a family or pair group. With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Figure 7—video 1). Moreover, while changes in the recording environment setup (e.g. lighting conditions, camera angles, etc.) do not affect the model performance, the model needs to be adjusted when family members or composition changes. In particular, the introduction of new members (e.g. newborns) requires assignment of new colored collar beads and collection of face images; thus, it is necessary that the recognition model of this marmoset family is retrained.

This system may be particularly useful when experimenters need to know which individual is performing a specific task for cognitive and behavioral experiments (Kangas et al., 2016; Kangas and Bergman, 2017; Marshall and Ridley, 2003). In these experiments, especially those involving long-term or continuous behavioral responses, experimenters often need to be present to record animal identity or to manually review video recordings after the experiment. This may distract the animals from their task and affect their behaviors, while also demanding sustained attention and a large time commitment for the experimenters. Our pipeline resolves this limitation by automatically detecting and recording marmoset identity throughout the experiment. A typical system requires the experimenter to only collect video clips from each marmoset, verify the automatically annotated marmoset faces and collar beads, and initiate the training of a family-specific face classification. This process takes approximately 20–30 hr of time. Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.

Future extensions can further improve the detection accuracy and the pipeline utility. First, the pipeline can be easily combined with additional programs to characterize the detailed facial features of the marmosets (Correia-Caeiro et al., 2022; Kawaguchi et al., 2023). Considering differences in their eye distance, mouth shape, and fur coloring pattern, in addition to the global facial structure, we can improve the identity prediction, particularly between two closely related marmosets with similar faces. Also, aside from the colored collar beads, experimenters can label marmosets’ identities using other visual markers, including color dyes on marmoset ear tufts. Using a more evident visual marker could be helpful for accurate identity prediction and avoid mislabeling. Moreover, it would be possible to integrate pose estimation tools, such as DeepLabCut and MarmoPose (Cheng et al., 2025; Lauer et al., 2022), with this marmoset recognition pipeline, thus allowing the estimation of the postures and behaviors of specific animals of interest. While we developed this pipeline and demonstrate its utility for common marmosets in laboratory captivity, there is no application restriction of this system. With appropriate training data and experimental design, our pipeline can be applied to other non-human primates in various settings such as lab housing, conservation fields, or even in the wild.

Materials and methods

Animals

Three adult common marmosets (C. jacchus, 1 father, 1 mother, 1 son, 6 years weighted at 464 g, 5 years at 566 g, and 3 years at 450 g, respectively) and two young common marmosets (two female twins, average weights 329 g and 338 g, aged 7–11 months) were involved in this study. The three adult marmosets belonged to the same family unit (parents: Adult1, Adult2, and their adult offspring: Adult3), while the young marmosets were twin siblings (Young1 and Young2) from another family. The marmosets are housed at The Neuro’s animal facility in family units in indoor enclosures. Housing cages included two sizes. The first size housed three adult marmosets with dimensions of 1.372 m length × 0.760 m width × 2.092 m height, and the second cage that housed the two young marmosets with their family has the dimensions of 1.065 m length × 1.067 m width × 2.092 m height. We placed a collar with a uniquely colored bead on each marmoset to allow consistent visual identification during the experiment. All experimental procedures were conducted under the Canadian Council on Animal Care guidelines, Standard Operating Procedures of marmosets, and Animal Use Protocol (AUP# 10000 and 10001) approved by The Neuro and McGill’s Animal Care Committee.

Video collection

Request a detailed protocol

Marmoset face dataset was collected by mounting a primate testing chair with an added camera component to the door of marmoset housing cages (Figure 9A and C). High-resolution videos were acquired using an industrial color/RGB camera (JIERUIWEITONG DF200-1080P). The industrial camera was selected specifically for its small size (camera box: 36 mm width × 36 mm height) and capability of close-distance recordings. High-resolution recordings (1920×1080 pixels, 30 frames per second) obtained with the small focal length lens (2.8 mm) enabled a wide field of camera view. This allowed the collection of full facial features and posture variability at close recording distances. The camera was adjusted for focus and placed approximately 10.50 cm from the housing cage door (Figure 9A and B). The camera was fixed in position using a transparent protective case, which was designed to prevent damage to the camera by the marmosets (Figure 9B).

Experimental setup of the camera system and primate chair for face detection and identification.

(A) Side view of the camera (green box) and primate chair, attached to the housing cage. (B) Frontal view of the camera fixed on the primate chair, enclosed within a protective cover. (C) A schematic illustration of the real-time facial image recording and automatic identification of the marmoset entering the primate chair.

Prior to the experiments, marmosets were acclimated to entering the primate testing chair space. This space minimized the chair surface reflection, reducing interference between true facial features and reflections during face detection and identification. Three adult marmosets from one family were recorded for 1 hr across each of the 5 recording days, with unrestricted voluntary access to the primate chair space. For the adult marmosets, the housing cage door was opened at the beginning of the recording session, allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 9C). Two young marmosets were briefly isolated and recorded separately for testing and improving the automatic face extraction program. We recorded them at two developmental time points, 7 months of age and 11 months of age (an additional time point at 16 months of age has been added for one marmoset to test the identification without collar). During video collection, a sliding panel and an in-cage box were positioned near the housing cage door to temporarily isolate individual marmosets from other family members. Individual isolation was kept brief (approximately 10 min) to prevent disturbance and potential stress due to family separation. Marmosets accessed the primate chair space through the housing cage door (11 cm width × 10.16 cm height).

Video preprocessing and bounding box annotation

Request a detailed protocol

The collected videos were preprocessed to identify which video clips included valid face and collar information of the marmoset. We defined valid face and collar features by whether the full-face details and colored bead appeared visible from the clips (Figure 10A, Step 1). To capture sufficient variability in postures, individuals, and lighting conditions, clips were sampled throughout the adult marmoset videos (approximately 5 hr) (Figure 10B). We used OpenCV to extract frames from the selected clips, and the extracted frames were reviewed to exclude those that were blurry (Bradski, 2000). Computer Vision Annotation Tool (CVAT) was applied to perform bounding box annotation of marmoset faces and collar colors from the extracted frames (Sekachev et al., 2020).

Workflow and design of the marmoset facial detection and identification model.

(A) The architecture of the real-time marmoset facial recognition program. (B) Marmoset face images from three camera angles. (C) Bounding boxes of marmoset faces (green box) and collars (pink box) were manually labeled in the training and validation datasets to train the multi-marmoset face classification model. (D) The schematic of the automatic face (blue box) and collar bead (cyan box) extraction model.

Marmoset face and identity dataset

Request a detailed protocol

To minimize image computations and data storage, we created a dataset of 2498 annotated images from the three adult marmosets. All images were manually annotated to label marmoset faces, individual identities, and the collar bead colors (Figure 10C). The annotated images were used for training models of multi-marmoset face classification and the automatic identity extraction, which can automatically detect, localize, and identify marmoset faces (Figure 10A, Steps 1–4). We created another dataset of two young marmosets at 7 months of age (total images = 502) for testing the automatic facial and identity extraction (Figure 10A, Steps 4–5). For both adult and young marmoset datasets, images were randomly divided into a training set and a validation set at a ratio of 8:2. Moreover, we evaluated the model performance and its generalization across developmental stages using new videos: (1) from the adult marmosets and (2) from the same young marmosets at 11 months, which were not involved during initial program training.

Marmoset face recognition pipeline

Request a detailed protocol

The marmoset facial detection and identification was primarily based on the YOLOv8 algorithm (Jocher et al., 2023; Varghese and Sambath, 2024). The architecture of the program consists of two main models: a multi-marmoset face classifier and an automatic facial and identity extraction (Figure 10A). In the initial step, we deployed multiple pre-trained YOLOv8 models (YOLOv8 nano, YOLOv8 small, and YOLOv8 medium) on the adult marmosets’ dataset (total images = 2498) to train the multi-marmoset face classifier (Lin et al., 2014). Training performance and detection accuracy of these models, generated by YOLO metrics as a CSV file, were evaluated to select the optimal model for subsequent training (see Results section). These pre-trained models share the same object detection backbone but differ in the number of parameters and computing power. Larger models, such as YOLOv8 medium, provide higher detection accuracy but require greater computational resources and longer inference time. In contrast, smaller models prioritize computational efficiency. The YOLOv8 nano model was selected and used for the training of all models in the program (Figure 10C). The adult marmosets’ dataset was also applied in the training of automatic facial and identity extraction. This trained automatic face extraction was utilized on the young marmoset dataset (total images = 502). Since young marmosets were recorded individually, we assigned the automatic annotation of marmoset faces and collar beads based on their recorded identity (Figure 10D). All automatic labeling was manually reviewed to remove mislabeled annotations. We used this dataset (total images = 449) to train the multi-marmoset face classifier for the young marmosets.

The marmoset facial recognition program was trained in the Anaconda virtual environment. The model training terminated at the early stopping criteria when no improvement was observed across 100 most recent epochs (i.e. training iterations). This tool involved training of three YOLOv8-based programs: (1) multi-marmoset face classifier for three adult marmosets was trained for 183 epochs; (2) automatic facial and identity extraction was trained for 124 epochs; (3) multi-marmoset face classifier for two young marmosets was trained for 245 epochs. The final models generated 2D bounding boxes, assigning identity and collar information of individual marmosets. The detection results of input videos were exported as text files, which represented the corresponding label identity, the bounding box size, and its location (Jocher et al., 2023; Redmon et al., 2015).

Model evaluation and analysis

The marmoset facial recognition program was evaluated using a range of performance evaluation metrics. The evaluation metrics included: precision, recall, mAP, validation DFL, and training time. We also computed the F1 score to select the best setup to be applied in the final real-time detection pipeline. In addition, we calculated the inter-individual face similarity between marmosets to test whether the biological face similarity between family members can explain the potential mislabeling of the face classifier programs.

Intersection over union

Request a detailed protocol

IoU calculates the amount of spatial overlapping between a predicted bounding box and the ground truth (i.e. manually labeled bounding box) (Jocher et al., 2023; Rezatofighi et al., 2019). It measures the localization accuracy and error amount between the predicted annotation and the manually labeled ground truth. A detection was considered correct if the IoU reached specific thresholds. In this study, we reported the mAP at the IoU thresholds from 0.5 to 0.95 (increments = 0.05). The IoU is calculated by the area of overlap (A∩B) and area of union (A∪B) of the predicted bounding box (A) and ground truth bounding box (B). If IoU = 1, the predicted bounding box is perfectly aligned with the true annotation. If IoU =0, there is no overlap between the two boxes.

IoU=A∩BA∪B

Precision, recall, and F1 score

Request a detailed protocol

Precision, recall, and F1 score are commonly used in the evaluation of classification models (Dehmer and Basak, 2012; Powers, 2008). Precision measures the overall proportion of correct face identifications over all positive results, indicating the accuracy of the detection model. Recall quantifies the proportion of the face identification that are correct compared with all actual positives, which reflects the ability of the model to predict correct marmoset face and collar beads. The F1 score is the harmonic mean of precision and recall, considering the false positives and false negatives in the model assessment.

Precision=TruepositiveTruepositive+Falsepositive
Recall=TruepositiveTruepositive+Falsenegative
F1 score=2×Precision×RecallPrecision+Recall

Mean average precision

Request a detailed protocol

mAP assesses how accurate the model detects objects and how well the model localizes the object in the image. In object detection algorithms, precision is the number of correctly detected objects, with recall being the number of true objects that are successfully detected. As the precision and recall are affected by different detection confidence thresholds, the performance is evaluated by the precision-recall curve, which plots the model precision (y-axis) against recall (x-axis) for a particular threshold. The average precision (AP) is the area under the precision-recall curve (Padilla et al., 2021). Before calculating the AP, the precision-recall pairs are interpolated to create a monotonically non-increasing precision-recall function, meaning that the interpolated precision (Pinterpolation) is the maximum precision (maxj≥iPrecisionj) at the recall level larger than or equal to Recalli. This interpolation assigns a highest precision value at each recall level, which smooths the precision-recall curve and reduces the impact of measurement fluctuations and noises to the evaluation of model performance.

Pinterpolation(Recalli)=maxj≥iPrecisionj

For each label class in the object detection model, the AP is the area under the curve of the interpolated precision-recall curve, calculated by summing the interpolated precision weighted by the incremental increase in the recall (Maxwell et al., 2021; Padilla et al., 2021).

AP=∑i=1N(Recalli−Recalli−1)Pinterpolation(Recalli)
N=total number of points on the precision−recall curve

The AP is extended further to calculate the mAP, which is the averaged AP values of different label classes in the object detection model (total class number = n). The APi is the mean AP calculated per label class at index i at multiple IoU levels (IoU = {0.50, 0.55, …, 0.95}). This calculation generates a more comprehensive evaluation with multiple label classes and their predicted localization accuracy of our marmoset facial recognition program.

mAP@50−95=1n∑i=1nAPi

Validation DFL and training time

Request a detailed protocol

The DFL and training time were also used to support the selection of the final marmoset facial recognition model. The DFL function measures the model’s performance to refine the bounding box predictions, based on the localization uncertainty of the bounding box annotations (Jocher et al., 2023). The predicted bounding box position is compared with the ground truth annotations, with a lower DFL value indicating a more precise bounding box prediction.

Training time was used the training time to assess the efficiency and the feasibility of the real-time marmoset face recognition program. The training time estimated the computational cost of the program, which provided additional information to select models with similar accuracy. The training time is calculated by the total time duration to train the model. For the number of training epochs, Nepoch, the total training time is calculated through multiplying Nepoch by the training time per epoch (tepoch):

Training time=Nepoch×tepoch

Inter-individual face similarity analysis

Request a detailed protocol

We calculated the inter-individual face similarity using customized Python scripts (see Data availability). The image feature embeddings, based on YOLOv8 algorithms, calculate the numerical representations (i.e. vectors) using a computer vision model, which encodes the semantic and visual maps of the image (Long Chai et al., 2023; Varghese and Sambath, 2024). In our marmoset facial recognition program, the intermediate layers included extensive feature maps of the visual details of the marmoset face images. By extracting and analyzing the image feature embeddings, we obtained various numerical representations of the marmoset faces, which allowed us to compare face shapes, distances between face features (eyes, mouth, nose, etc.), and the fur patterns of different individuals. We assessed the inter-individual face similarities for two family groups: (1) within the adult marmoset family (parents and adult offspring) and (2) between the two young marmosets (twin siblings). For each individual marmoset, we selected 10 images with similar head orientation, lighting conditions, and camera angles, and cropped them to only include the face regions. As the marmoset face images passed through the YOLO convolutional layers, the images were outputted into a feature map (Aghdam and Heravi, 2017; Dumoulin and Visin, 2016; Long Chai et al., 2023; Redmon et al., 2015). For one marmoset face image, the feature map (Fi) is encoded in the number of inputs (ninput) and the feature map shapes (Aghdam and Heravi, 2017; Dumoulin and Visin, 2016).

Fi(h,w,c)=ninput×height×width×channel

We extracted face embeddings from the face classifiers of the adult marmosets’ family (parents and son) and young twin marmosets from another family. To generate a single vector (i.e. embedding, ei) representing the marmoset face feature per channel, we averaged across the spatial dimension (i.e. height and width of the image) per channel for one marmoset face image (Lee et al., 2025). The extracted embedding (ei) was averaged across the 10 images for each individual marmoset, which allowed us to have an averaged embedding vector (emarmoset) and feature maps to calculate the inter-individual face similarities. We performed L2 normalization on this embedding (emarmoset) to create a unit vector (enorm) for inter-individual comparison, using scikit-learn preprocessing function (Pedregosa et al., 2011).

ei=1height1width∑h=1height∑w=1widrthFi(h,w,c)
emarmoset=110∑i=110ei
enorm=emarmoset‖emarmoset‖,where‖emarmoset‖=∑emarmoset2

As such, with the normalized embedding vectors from the marmosets in this study, we calculated the inter-individual face similarities between marmoset pairs using cosine similarity (Li and Han, 2013; Nguyen and Bai, 2011; Thongtan and Phienthrakul, 2019). Moreover, we also calculated the Euclidean distances between extracted embeddings from two marmosets (Jozwik et al., 2022; Sugase-Miyamoto et al., 2014). The calculation equations were described using the normalized embedding vectors of a marmoset pair (enorm – 1 and enorm – 2) in scikit-learn library (Merchant et al., 2023; Pedregosa et al., 2011):

Cosine similarity=enorm−1⋅enorm−2||enorm−1|| ||enorm−2||
Euclidean distance=∑i=1n(enorm−1−enorm−2)2

We plotted a violin plot of cosine similarity and Euclidean distance to interpret the inter-individual facial similarities calculated from the marmoset pairs’ face embeddings. In addition, the similarity values were standardized using within-model z-scores.

Statistical tests were performed only within the face classifier of each marmoset family, as embedding spaces may vary in scaling, learned features, and baseline metrics making cross-model comparison of inter-individual face similarity unreliable (Bollegala, 2017). To test whether the learned facial embeddings between different marmosets are related to the biological facial feature similarities, we performed the pairwise Welch’s t-test and the Cohen’s d to examine the significance and effect size. Within each family unit, facial similarity was compared between two types of family relationships (e.g. mother-son vs. father-son).

Real-time face recognition program

Request a detailed protocol

We performed the real-time marmoset face detection and identification using the customized Python scripts, based on the trained weights from the final multi-marmoset facial classification model in YOLOv8. Live videos were collected using the color (RGB) camera and simultaneously processed frame-by-frame to detect the marmoset identities. For each detected bounding box, the scripts returned a corresponding label of marmoset face and collar bead color. We assigned the detected collar beads as the corresponding marmoset identity with a higher weight, which improved the detection confidence across frames.

To ensure the program accuracy and efficiency, we temporally stored the detection results of the most recent 30 frames (approximately 1 s). The most frequently detected marmoset identity within this time window was exported as the current output. We wrote and updated the detection results continuously into an output JSON file, enabling real-time identity reading of the marmoset who entered the primate chair most recently.

Appendix 1

Appendix 1—table 1
Model performance evaluation of the validation dataset in the adult marmoset recognition model.

Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.

Label classesPrecisionRecallmAP50mAP@50–95
All0.9360.9550.9800.713
Adult10.8290.9450.9570.841
Adult20.9260.9830.9910.787
Adult30.9510.9980.9920.868
collar_Adult10.9680.9660.9880.603
collar_Adult20.9780.9720.9900.665
collar_Adult30.9620.8660.9590.515
Appendix 1—table 2
Model performance evaluation of the validation dataset in the automatic face and identity extraction model.

Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.

Label classesPrecisionRecallmAP50mAP@50–95
All0.9380.9680.9840.715
Face0.9190.9860.9900.831
Collar0.9570.9500.9780.600
Appendix 1—table 3
Model performance evaluation of the validation dataset in the young marmoset recognition model.

Comparison of precision, recall, mAP@50, and mAP@50–95 were computed for all the detected label classes in this model.

Label classesPrecisionRecallmAP50mAP@50–95
All0.9790.9750.9940.890
Young11.0000.9350.9930.950
collar_Young10.9981.0000.9950.808
Young20.9581.0000.9950.947
collar_Young20.9630.9640.9930.855

Data availability

All code in this paper is publicly available at GitHub: https://github.com/Jy-Yang-bot/real-time-marmoset-recognition (copy archived at Yang, 2026). This repository contains the scripts for pre-processing the images, the main facial recognition pipeline, and the analytic tools. The program generalizability was tested on additional video obtained from publicly available online sources, including a YouTube video "Common Marmoset", uploaded by Billabong Zoo, Koala & Wildlife Park on 2 June 2020 (https://www.youtube.com/watch?v=lDJNk4rjqqQ). As per McGill policy on sharing and use of images of animals involved in Research, Teaching and Testing, we cannot upload the raw videos recorded within our animal facilities into a domain-specific public archive or generic repositories. However, those raw videos can be made available upon request by reaching out to the corresponding authors (Jiayue Yang and Justine Cléry).

References

    1. Bradski G
    (2000)
    The opencv library
    Dr. Dobb’s Journal of Software Tools 25:120–125.
  1. Book
    1. Li B
    2. Han L
    (2013) Distance weighted cosine similarity measure for text classification
    In: Yin H, Tang K, Gao Y, Klawonn F, Lee M, Weise T, Li Bin, Yao X, editors. Intelligent Data Engineering and Automated Learning – IDEAL 2013, Lecture Notes in Computer Science. Berlin Heidelberg: Springer. pp. 611–618.
    https://doi.org/10.1007/978-3-642-41278-3_74
  2. Book
    1. Lin TY
    2. Maire M
    3. Belongie S
    4. Bourdev L
    5. Girshick R
    6. Hays J
    7. Perona P
    8. Ramanan D
    9. Zitnick CL
    10. Dollár P
    (2014) Microsoft COCO: common objects in context
    In: Fleet D, Pajdla T, Schiele B, Tuytelaars T, editors. Computer Vision – ECCV 2014. Springer. pp. 740–755.
    https://doi.org/10.1007/978-3-319-10602-1_48
  3. Book
    1. Merchant FA
    2. Shah SK
    3. Castleman KR
    (2023) Object measurement
    In: Merchant FA, Castleman KR, editors. Microscope Image Processing. Elsevier. pp. 153–175.
    https://doi.org/10.1016/B978-0-12-821049-9.00017-4
  4. Book
    1. Nguyen HV
    2. Bai L
    (2011) Cosine similarity metric learning for face verification
    In: Kimmel R, Klette R, Sugimoto A, editors. Computer Vision – ACCV 2010, Lecture Notes in Computer Science. Springer. pp. 709–720.
    https://doi.org/10.1007/978-3-642-19309-5_55
    1. Pedregosa F
    2. Varoquaux G
    3. Gramfort A
    4. Michel V
    5. Thirion B
    6. Grisel O
    7. Blondel M
    8. Prettenhofer P
    9. Weiss R
    10. Dubourg V
    11. Vanderplas J
    12. Passos A
    13. Cournapeau D
    14. Brucher M
    15. Perrot M
    16. Duchesnay E
    (2011)
    Scikit-learn: machine learning in python
    Journal of Machine Learning Research 12:2825–2830.
  5. Conference
    1. Thongtan T
    2. Phienthrakul T
    (2019) Sentiment Classification Using Document Embeddings Trained with Cosine Similarity
    Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. pp. 407–414.
    https://doi.org/10.18653/v1/P19-2057
  6. Conference
    1. Varghese R
    2. Sambath M
    (2024) YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness
    2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS). pp. 1–6.
    https://doi.org/10.1109/ADICS58448.2024.10533619
  7. Conference
    1. Yu L
    2. Chen X
    3. Zhou S
    (2018) Research of Image Main Objects Detection Algorithm Based on Deep Learning
    2018 IEEE 3rd International Conference on Image, Vision and Computing (ICIVC). pp. 70–75.
    https://doi.org/10.1109/ICIVC.2018.8492803

Article and author information

Author details

  1. Jiayue Yang

    1. The Neuro, Department of Neurology and Neurosurgery, McGill University, Montreal, Canada
    2. Integrated Program in Neuroscience, McGill University, Montreal, Canada
    Contribution
    Conceptualization, Resources, Data curation, Software, Formal analysis, Funding acquisition, Validation, Investigation, Visualization, Methodology, Writing – original draft, Writing – review and editing, Data collection
    For correspondence
    jiayue.yang@mail.mcgill.ca
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0009-0003-3727-9586
  2. James Wang

    1. The Neuro, Department of Neurology and Neurosurgery, McGill University, Montreal, Canada
    2. Integrated Program in Neuroscience, McGill University, Montreal, Canada
    Contribution
    Investigation, Methodology, Writing – review and editing, Data collection
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0009-0002-6656-7176
  3. Justine Cléry

    1. The Neuro, Department of Neurology and Neurosurgery, McGill University, Montreal, Canada
    2. McConnell Brain Imaging Centre, The Neuro, Montreal Neurological Institute and Hospital, McGill University, Montreal, Canada
    3. Azrieli Centre for Autism Research, The Neuro, Montreal, Canada
    Contribution
    Conceptualization, Resources, Supervision, Funding acquisition, Validation, Investigation, Methodology, Project administration, Writing – review and editing
    For correspondence
    justine.clery@mcgill.ca
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0003-1020-1845

Funding

New Frontiers in Research Fund (NFRFT-2022-00051)

  • Justine Cléry

Fonds de recherche du Québec

https://doi.org/10.69777/347426
  • Justine Cléry

Fonds de recherche du Québec (358082)

  • Justine Cléry

Fonds de recherche du Québec

https://doi.org/10.69777/2003934
  • Jiayue Yang

The funders had no role in study design, data collection and interpretation, or the decision to submit the work for publication.

Acknowledgements

We would like to thank our animal health technicians D Hau-Aquino, V Comtois, C Hunt; and veterinarians F Chaurand, and J Hutta for animal health care and support. We are thankful to M Gacoin and T Cook for helpful discussions and feedback on the visualization and format of the manuscript. We thank J Smith and C O'Hare-Freire for the design and construction of a transparent protective case for the camera. We acknowledge the support of the Government of Canada’s New Frontiers in Research Fund (NFRF) (NFRFT-2022-00051) and by the Fonds de Recherche du Québec-Santé (FRQS) (#347426, #358082, and #2003934). Ces travaux ont bénéficié d’un octroi des fonds Nouvelles frontières en recherche du gouvernement du Canada (NFRFT-2022-00051) et du Fonds de recherche du Québec-Santé (FRQS, #347426, #358082, et #2003934).

Ethics

All experimental procedures were conducted under the Canadian Council on Animal Care guidelines, Standard Operating Procedures of marmosets, and Animal Use Protocol (AUP# 10000 and 10001) approved by The Neuro and McGill's Animal Care Committee.

Version history

  1. Preprint posted:
  2. Sent for peer review:
  3. Reviewed Preprint version 1:
  4. Reviewed Preprint version 2:
  5. Version of Record published:

Cite all versions

You can cite all versions using the DOI https://doi.org/10.7554/eLife.110932. This DOI represents all versions, and will always resolve to the latest one.

Copyright

© 2026, Yang et al.

This article is distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use and redistribution provided that the original author and source are credited.

Metrics

  • 566
    views
  • 34
    downloads
  • 0
    citations

Views, downloads and citations are aggregated across all versions of this paper published by eLife.

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Jiayue Yang
  2. James Wang
  3. Justine Cléry
(2026)
A real-time, multi-animal model for automatic face detection and identification of freely moving common marmosets based on YOLOv8 algorithms
eLife 15:RP110932.
https://doi.org/10.7554/eLife.110932.3

Share this article

https://doi.org/10.7554/eLife.110932