Abstract
An objective assessment of a mentor’s behavioral skills is needed to assess the effectiveness of mentor training interventions in academic settings. The Mentor Behavioral Interaction (MBI) Rubric is a newly developed, content-valid, observational measure of a mentor’s behavioral skill during single-episode interactions with a mentee. The purpose of this study was to assess the inter-rater reliability (IRR) of the MBI Rubric when used to assess video-recorded mentor-mentee interactions. Three of a pool of four faculty raters with expertise in mentor training synchronously rated 26 videos of mentor-mentee interactions using structured guidelines. The MBI Rubric includes six items (Part 1), each with ratings on a 3- or 4-point scale, and ten yes/no items (Part 2) that characterize the content of the interaction. After initial individual ratings were completed, the three raters met, reviewed disagreements, and reached decisions about final item scores by either consensus or majority vote. Mean total Part 1 scores ranged between 1.42–2.69. IRRs ranged from good (Part 1 IRR=0.67) to excellent (Part 2 IRR=0.83). No training effects were observed, with no decrease (i.e., showing less variability) in inter-rater standard deviations over time. Rater effects in initial individual scoring were observed, with a significant difference between one vs. the other three raters on Part 1 individual scores, with no effects for Part 2 scores. Raters tended to score lower on initial individual scores than the final score for both Part 1 and 2. The MBI Rubric is the first observational measure to assess single episodes of video-recorded mentor-mentee interactions and has demonstrated content validity, and now inter-rater reliability. It may be used in parallel with other instruments to measure the efficacy of mentor training. Limitations include possible ceiling effects, and resource-intensive administration in terms of rater expertise and time. Future work will assess the responsiveness of the Rubric to change in mentor skill and construct validity.
Introduction & Literature Review
Outcomes such as increased research productivity, leadership skills, career success, career satisfaction, and career commitment have been correlated with effective mentoring (Hafsteinsdóttir et al., 2017; Libby et al., 2016; National Academies of Sciences, Engineering, and Medicine, 2019; Sambunjak et al., 2006). The National Academies of Science, Engineering and Medicine (2019) strongly recommended that institutions transition from “a culture of ad hoc mentorship…toward one of intentional, inclusive and effective mentorship” (p. 7). As a result of these observations and recommendations, researchers have increasingly focused on improving faculty mentoring skills (Pfund et al., 2014; Sood et al., 2020). Reliable and valid self-report measures, such as the Mentoring Competency Assessment (Fleming et al., 2013) are available to test the efficacy of such proposed interventions, but until recently there have been no objective measures to complement self-administered scales. The new Mentor Behavioral Interaction (MBI) Rubric has demonstrated content validity as a measure of a mentor’s behavioral skill with a mentee in single-episode interactions (Tigges et al., 2022). The purpose of this study was to assess the inter-rater reliability (IRR) of the MBI Rubric when used to assess video-recorded mentor-mentee interactions.
Methods
Sample
The sample for this analysis was 26 unique mentor-mentee pairs and their video-recorded online Zoom meetings gathered as part of data collection for a larger study of mentor training interventions at multiple Southwest and Mountain West institutions, including the University of New Mexico (UNM) Central and Health Science campuses, Arizona State University, Oklahoma University Health Science Center, and Mountain West Clinical Translational Research Infrastructure Network institutions. The meetings ranged in length from 34 to 114 minutes (M = 55.5; SD = 15.0). The study was approved by the UNM Health Sciences Center Institutional Review Board (HRPO 18–261).
Instrument
The MBI Rubric was used to score the mentor’s behavioral skills during single-episode, video-recorded mentor-mentee interactions. The development of the Rubric is largely based on Fleming et al.’s (2013) Mentoring Competency Assessment, and its content validity has been described elsewhere (Tigges et al., 2022). Part 1 of the MBI Rubric forms the primary measure of mentor behavior and includes six items, each with ratings on a 3- or 4-point scale (see Table 1). Higher scores indicate higher mentor performance. Three of the six items (constructive feedback, acknowledging professional contributions, and helping mentee to network) had the lowest scores (0) that were negative behaviors and also had a not observed (1) scoring option, followed by 2 (mid-performance), and 3 (highest performance). The remaining three items (active listening, reflective listening, and motivating mentees) did not include a not observed option in scoring because the lowest possible score (1) was that the behavior was not observed, followed by 2 (mid-performance) and 3 for (highest performance). Part 2 of the Rubric contains 10 yes/no items that characterize the content of the interaction and are used for descriptive purposes only. An example of a Part 2 item is “Discusses integration of work with personal life.” Total scores for each part were calculated by summing individual item scores and dividing by the number of items. Possible total scores ranged from 0.5 to 3 for Part 1, and 0 to 10 for Part 2.
Table 1.
Part 1 Final Score Descriptive Statistics
| Part 1 Item | N | M | SD | Median | Minimum | Maximum | # Not Observed (Score=1) |
|---|---|---|---|---|---|---|---|
| 1. Actively listens b | 26 | 2.42 | 0.81 | 3 | 1 | 3 | 5 |
| 2. Uses reflective listening b | 26 | 1.42 | 0.64 | 1 | 1 | 3 | 17 |
| 3. Provides constructive feedback a | 26 | 1.65 | 0.69 | 2 | 1 | 3 | 12 |
| 4. Motivates mentee b | 26 | 2.69 | 0.55 | 3 | 1 | 3 | 1 |
| 5. Acknowledges professional contributions a | 26 | 2.15 | 0.67 | 2 | 1 | 3 | 4 |
| 6. Helps mentee to network a | 26 | 2.27 | 0.72 | 2 | 1 | 3 | 4 |
| Total Part 1 Score | 26 | 2.10 | 0.42 | 2 | 1 | 3 | n/a |
Score options: 0=lowest performance (negative behavior); 1=no observed behavior; 2=middle performance; 3=highest performance.
Score options: 1=lowest performance (no observed behavior); 2=middle performance; 3=highest performance.
Prior to the actual use of the Rubric, investigators developed an MBI Rubric Scoring Guidelines Manual with scoring instructions for each item. The initial manual was tested on sample videos and modified as needed. In addition, during the formal review of the initial set of videos, investigators continued to make minor edits and clarifications to the manual, although the Rubric and its scoring categories were never changed.
Procedures
Three out of a pool of four faculty raters with expertise in mentor training met via Zoom to synchronously rate each video using a REDCap™ scoring sheet (Harris et al., 2009). Three raters, rather than four, were used for scoring to allow for flexibility in scheduling; final raters were chosen based on availability to participate in scoring. After initial individual ratings were completed, the three raters reviewed disagreements in scoring for each item, discussed the rationale for their individually assigned score, and reached consensus or majority decisions about the assignment of final item scores.
Analysis
Inter-rater reliability (IRR) is a statistical measure used to assess the consistency or agreement among different raters when evaluating the same set of data, in this case, the mentor’s behavior in video-recorded interactions. IRR ranges from 0 to 1, and is a measure of how much of the variance in observed scores is due to variance in true scores after measure error variance between raters is removed. In this study, intraclass correlation coefficients (ICC) were used as the measure of IRR using two-way random effects models with random videos and random raters (generalizing to all potential raters) for absolute agreement on the average score of multiple raters (Hallgren, 2012; Koo & Li, 2016; Shrout & Fleiss, 1979). The design was not completely crossed as only three of four raters scored each video. Higher ICCs indicate higher IRR, with a value of 0 indicating only random agreement to 1 indicating perfect agreement. Values between .60 and .74 may be considered good and between .75 and 1.0 as excellent (Hallgren, 2012).
Results
Table 1 shows the descriptive characteristics for the six Part 1 items that form the main score for the MBI Rubric. Mean total Part 1 final scores ranged between 1.42 out of a possible 3.0, for the reflective listening item, and 2.69, for the motivates mentee item. Two items, uses reflective listening (17 out of 26), and provides constructive feedback (12 out of 26), were the most commonly not observed (score = 1). Reflective listening was defined as “verbally reflecting
Average score calculated by dividing sum of scores by number of scores. Possible total score = 0.5 to 3.0. back and paraphrasing what the mentee has said to ensure understanding and creates a feeling of respect and value.” Constructive feedback was defined as “evaluative or corrective information from the mentor to the mentee about a specific action, event, or process that has already occurred that goes beyond simply listing the achievement.” No negative behaviors were observed on the three items, denoted by superscript “a” in the table, that had possible lowest scores of zero: provides constructive feedback, acknowledges professional contributions, and helps mentee to network. With the exceptions of “uses reflective listening” and “provides constructive feedback”, mean scores were between middle and highest performance (scores of 2–3). Two of the six items had median scores of 3: active listening and motivates mentee.
IRRs ranged from good (Part 1 IRR=0.67) to excellent (Part 2 IRR=0.83). Figure 1 shows that neither Part 1 nor Part 2 standard deviations were correlated with video order, rs = −.04, p = .86; rs = −.26, p = .20, respectively. In other words, in this small sample, there were no training effects observed (i.e., no decrease in inter-rater standard deviation, or less variability, in scoring) over time as raters gained experience using the MBI Rubric and accompanying training manual.
Figure 1.

Trends in Inter-Rater Standard Deviation (SD) Scores Over Time (Videos 1–26)
For each MBI Rubric item, two types of scores were generated: Three initial individual rater scores based on viewing of the video-recorded interaction, and one final score assigned by either consensus or majority vote after raters discussed their individual scoring. Significant rater effects on initial scoring were observed, with significant differences between one versus the three other raters on Part 1 total scores (see Table 2). Rater 1 tended to score lower than the other three raters. No significant rater effects on initial scoring were observed for Part 2 scores (yes/no scoring). Table 3 shows changes in scores from initial individual rater scores to final consensus or majority score for both Part 1 and Part 2. On Part 1, three of the four raters had initial scores that were lower than the final score. There was a significant rater effect for initial to final changes in Part 1 scores, largely because of a difference between Rater 1, who had a significant change in scores, when compared to Rater 4, who had the smallest change in scores (Table 3). In contrast, for Part 2, there were no significant rater effects in initial to final score change, although Rater 2’s score significantly changed the most. Three of the four raters had initial scores that were lower than final Part 2 scores, i.e., may not have observed behaviors that other scorers did when completing the yes/no scoring. Rater 1 had the least change in Part 2 scoring.
Table 2.
Assessment of Rater Effects: Part 1 and Part 2 Mean Scores and Standard Errors by Rater
| Rater | Part 1 Score (Range 0.5–3) |
Part 2 Score (Range 0–10; Yes/No) |
||
|---|---|---|---|---|
| M | SE | M | SE | |
| 1 | 1.76 | .09 | 8.23 | .42 |
| 2 | 2.08 | .09 | 7.20 | .41 |
| 3 | 2.09 | .11 | 7.35 | .47 |
| 4 | 2.19 | .09 | 7.62 | .41 |
| Total | 2.10 | 8.08 | ||
| Part 1 Rater Effect: p = .001 | Part 2 Rater Effect: p = .060 | |||
| Rater 1 vs. Rater 2, p = .003 | Rater 1 vs. Rater 2, p = .057 | |||
| Rater 1 vs. Rater 3, p = .043 | ||||
| Rater 1 vs. Rater 4, p = <.001 | ||||
Note: Dependent variable is Part 1 or Part 2 score by each rater. Mixed model with fixed rater effect and with random video and residual errors. Rater effect statistical testing adjusted for multiple comparisons
Table 3.
Initial Individual Rater Score Minus Final Consensus/Majority Score: Changes in Part 1 and Part 2 Scores After Discussion
| Rater | Change in Part 1 Score | Change in Part 2 Score | ||||
|---|---|---|---|---|---|---|
| M | SE | 95% CI | M | SE | 95% CI | |
| 1 | −1.65 | .43 | [−2.51, −0.079] | 0.16 | .28 | [−0.41, 0.72] |
| 2 | −0.23 | .41 | [−1.05, 0.60] | −0.83 | .27 | [−1.37, −0.29] |
| 3 | −0.43 | .51 | [−1.46, 0.60] | −0.66 | .34 | [−1.34, 0.01] |
| 4 | 0.09 | .42 | [−0.75, 0.94] | −0.51 | .28 | [−1.06, 0.04] |
| Total | −0.55 | .23 | [−1.02, −0.75] | −0.46 | .15 | [−0.77, −0.14] |
| Part 1 Rater Effect: p = .032 | Part 2 Rater Effect: p = .090 | |||||
| Rater 1 vs. Rater 2, p = .117 | Rater 1 vs. Rater 2, p = .068 | |||||
| Rater 1 vs. Rater 3, p = .371 | ||||||
| Rater 1 vs. Rater 4, p = .033 | ||||||
Note: Negative score indicates initial score lower than final consensus/majority score and reflects that the individual scorer may not have observed behaviors that other scorers did when completing the scoring. Dependent variable is Part 1 or Part 2 score by each rater. Mixed model with fixed rater effect and with random video and residual errors. Rater effect statistical testing adjusted for multiple comparisons
Discussion
This study has established the IRR of the newly developed MBI Rubric, the first observational measure to assess single episodes of video-recorded mentor behavior in interaction with a mentee. With this group of raters, who also developed the MBI Rubric, raters did not get more consistent over time, although the sample size was small. As might be expected, there were rater effects for the Part 1 measure, with one rater assigning significantly initial lower scores than the other three and demonstrating the most change from the initial individual score to a higher final consensus/majority score. Despite these differences, the MBI Rubric demonstrated acceptable IRR, particularly for a new measure. The consistency in IRR over time would be important to test over a longer period of time with additional videos and with additional raters who are not as familiar with the Rubric at the start of scoring. It is recommended that prior to using the MBI Rubric for actual research- or teaching-related scoring, new raters familiarize themselves with the Rubric scoring guidelines, as well as use sample videos to train themselves.
This initial study using this new measure also provides insight into video-recorded patterns of behavior for this group of mentors and suggests skills to more strongly emphasize in mentor training programs. First, the use of reflective listening was, by far, the weakest mentoring behavior observed, with over half of the 26 mentors not demonstrating this skill and not reflecting back either the mentee’s verbal statements or their emotional state. In particular, even in circumstances when mentees were near tears, mentors seemed reluctant to directly address the mentee’s emotional state. Although almost all mentors provide a great deal of advice, only about half of the mentors were observed providing constructive feedback, defined in this study as evaluative or corrective information about a specific action, event, or process that has already occurred. That is, few mentors were observed providing specific feedback about written materials such as manuscripts, methods proposals, or individual development plans that mentees had provided prior to the meeting. In contrast, to some degree, more than half of the mentors actively listened to their mentees, including asking clarifying questions; motivated their mentees through words of persuasion or praise, as well as expressing excitement or conveying positivity about the mentee’s work, or offering the mentor’s own optimistic examples as encouragement; and actively helped their mentees to network through either sharing information or providing relationship advice. Whether these observations are typical of mentor-mentee relationships over time, or unique to any single observation of a mentor-mentee meeting, remains to be seen.
Tigges et al. (2022) acknowledged several limitations of the MBI Rubric, including that it focuses only on the mentor’s behavior, evaluates a single interaction with a mentee, and does not obtain any concurrent ratings of cognitive processes (such as perceptions of motivation, acknowledgement, or emotional support). This analysis highlights two additional limits. First, the MBI Rubric was developed to measure improvement in mentors’ behaviors after mentor training. Yet scores on these first 26 videos are quite high, raising concerns about ceiling effects and the inability to detect any change after an intervention. It is possible that mentors and mentees behave differently when they know they are being observed. Second, the methods used to score these videos are resource-intensive, both in terms of the raters involved and actual time. To capture the nuances of mentoring, senior faculty who are experienced mentors were chosen as raters. Each rating session for a single video took approximately 1–1/2 hours for all three faculty raters, a total of 4–1/2 person hours per video. Initial experience from this study suggests that three raters may be ideal for capturing all the detailed content in an approximately one-hour interaction. Additional assessments of reliability will be needed to see if less senior raters and fewer than three raters can provide equally reliable results. Future work will also include assessments of the MBI Rubric’s responsiveness to change in mentor skill and further tests of construct validity. In addition, this study was conducted during the SARS-CoV-2 pandemic, when video interactions between mentors and mentees became the norm. Face-to-face interactions, if they resume, will need to be examined separately in the future.
Conclusion
The MBI Rubric is a highly novel, content-valid, reliable observational measure of the quality of a faculty mentor’s behavioral interaction with their faculty mentee in single episodes of interaction. The observed MBI Rubric may be used in parallel with other self-reported instruments and objective measures of career success to measure the efficacy of mentoring and mentor training.
Acknowledgment
Funded by the NIH/NIGMS U01GM132175 (Sood, A.) and 2U54GM104944 (Sy, F.); HRSA grant 1 D34HP45723-01-00 (PI Romero-Leggott); and NIH/NCATS UL1 TR001449 (Pandhi, N./Campen, M.).
References
- Fleming M, House S, Hanson VS, Yu L, Garbutt J, McGee R, Kroenke K, Abedin Z, & Rubio DM (2013). The mentoring competency assessment: Validation of a new instrument to evaluate skills of research mentors. Academic Medicine, 88, 1002–1008. https://dx.doi.org/10.1097%-2FACM.0b013e318295e298 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hafsteinsdóttir TB, van der Zwaag AM, & Schuurmans MJ (2017). Leadership mentoring in nursing research, career development and scholarly productivity: A systematic review. International Journal of Nursing Studies, 75, 21–34. 10.1016/j.ijnurstu.2017.07.004 [DOI] [PubMed] [Google Scholar]
- Hallgren KA (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23–34. 10.20982/tqmp.08.1.p023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Harris PA, Taylor R, Thielke R, Payne J, Gonzalez N, & Conde JG (2009). Research electronic data capture (REDCap)--A metadata-driven methodology and workflow process for providing translational research informatics support. Journal of Biomedical Informatics, 42(2), 377–381. 10.1016/j.jbi.2008.08.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Koo TK, & Li MY (2016). A guideline for selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15, 155–163. 10.1016/j.jcm.2016.02.012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Libby AM, Hosokawa PW, Fairclough DL, Prochazka AV, Jones PJ, & Ginde AA (2016). Grant success for early-career faculty in patient-oriented research: Differences-in-differences evaluation of an interdisciplinary mentored research training program. Academic Medicine, 91, 1666–1675. 10.1097/ACM.0000000000001263 [DOI] [PMC free article] [PubMed] [Google Scholar]
- National Academies of Sciences, Engineering, and Medicine. (2019). The science of effective mentorship in STEMM Washington, DC, United States: The National Academies Press. 10.17226/25568 [DOI] [PubMed] [Google Scholar]
- Pfund C, House SC, Asquith P, Fleming MF, Buhr KA, Burnham EL, Gilmore JME, Huskins WC, McGee R, Schurr K, Shapiro ED, Spencer KC, & Sorkness CA (2014). Training mentors of clinical and translational research scholars: A randomized controlled trial. Academic Medicine, 89, 774–782. 10.1097/ACM.0000000000000218 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Sambunjak D, Straus SE, & Marusic A (2006). Mentoring in academic medicine: A systematic review. Journal of the American Medical Association, 296(9), 1103–1115. 10.1001/jama.296.9.1103 [DOI] [PubMed] [Google Scholar]
- Shrout PE, & Fleiss JL (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. 10.1037/0033-2909.86.2.420 [DOI] [PubMed] [Google Scholar]
- Sood A, Qualls C, Tigges B, Wilson B, & Helitzer D (2020). Effectiveness of a faculty mentor development program for scholarship at an academic health center. Journal of Continuing Education in the Health Professions, 40(1), 58–65. 10.1097/CEH.0000000000000276 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Tigges B, Sood A, Mickel N, Dominguez N, Helitzer D (2022). Development and content validity testing of the mentor behavioral interaction rubric. Chronicles of Mentoring and Coaching, 6(Spec Iss 15), 630–636. [PMC free article] [PubMed] [Google Scholar]
