Methods: This was a prospective educational study with comparison with historical controls (reference cohort). At a temporal bone dissection course, eighteen participants performed structured self-assessment during three hours of VR simulation training of mastoidectomy before proceeding to cadaver dissection/surgery (intervention cohort). At a previous course, eighteen participants received similar VR simulation training but without the structured self-assessment (reference cohort). Final products from VR simulation and cadaveric dissection were video-recorded and assessed by two blinded raters using a 19-point modified Welling Scale.
Results: The intervention cohort completed fewer procedures (average 4.2) during VR simulation training than the reference cohort (average 5.7). Nevertheless, the intervention cohort achieved a significantly higher average dissection score both in VR simulation (11.1 points, 95% CI [10.6–11.5]) and subsequent cadaveric dissection (11.8 points, 95% CI [10.7–12.8]) compared with the reference cohort who scored 9.1 points (95% CI [8.7–9.5]) during VR simulation and 5.8 points (95% CI [4.8–6.8]) during cadaveric dissection.
Conclusion: Structured self-assessment is a valuable learning support during self-directed VR simulation training of mastoidectomy and the positive effect on performance transfers to subsequent cadaveric dissection performance.
]]>Summary of background: Simulation-based training is increasingly used in surgical education. However, it is important to determine which level of competency trainees must reach during simulation-based training before operating on patients. Therefore, pass/fail standards must be established using systematic, transparent, and valid methods.
Methods: Systematic literature search was done in four databases (Ovid MEDLINE, Embase, Web of Science, and Cochrane Library). Original studies investigating simulation-based assessment of surgical procedures with application of a standard setting were included. Quality of evidence was appraised using GRADE.
Results: Of 24,299 studies identified by searches, 232 studies met the inclusion criteria. Publications using already established standard settings were excluded (N = 70), resulting in 162 original studies included in the final analyses. Most studies described how the standard setting was determined (N = 147, 91%) and most used the mean or median performance score of experienced surgeons (n = 65, 40%) for standard setting. We found considerable differences across most of the studies regarding study design, set-up, and expert level classification. The studies were appraised as having low and moderate evidence.
Conclusion: Surgical education is shifting towards competency-based education, and simulation-based training is increasingly used for acquiring skills and assessment. Most studies consider and describe how standard settings are established using more or less structured methods but for current and future educational programs, a critical approach is needed so that the learners receive a fair, valid and reliable assessment.
]]>Methods: Prospective, single-arm trial. Twenty-four novice medical students completed a pre-training CI inserting test on a commercially available pre-drilled 3D-printed temporal bone. A training program of 18 VR simulation CI procedures was completed in the Visual Ear Simulator over four sessions. Finally, a post-training test similar to the pre-training test was completed. Two blinded experts rated performances using the validated Cochlear Implant Surgery Assessment Tool (CISAT). Performance scores were analyzed using linear mixed models.
Results: Learning curves were highly individual with primary performance improvement initially, and small but steady improvements throughout the 18 procedures. CI VR simulation performance improved 33% (p < 0.001). Insertion performance on a 3D-printed temporal bone improved 21% (p < 0.001), demonstrating skills transfer.
Discussion: VR SBT of CI surgery improves novices’ performance. It is useful for introducing the procedure and acquiring basic skills. CI surgery training should pivot on objective performance assessment for reaching pre-defined competency before cadaver – or real-life surgery. Simulation-based training provides a structured and safe learning environment for initial training.
Conclusion: CI surgery skills improve from VR SBT, which can be used to learn the fundamentals of CI surgery.
]]>Method: Twenty-four medical students were randomised in two groups and performed 15 mastoidectomies on a distributed virtual reality simulator as practice. The intervention group received additional summative metrics-based feedback; the control group followed standard instructions. Two to three months after training, participants performed a retention test without learning supports.
Results: The intervention group had a better final-product score (mean difference = 1.0 points; p = 0.001) and metrics-based score (mean difference = 12.7; p < 0.001). At retention, the metrics-based score for the intervention group remained superior (mean difference = 6.9 per cent; p = 0.02). Also at the retention, cognitive load was higher in the intervention group (mean difference = 10.0 per cent; p < 0.001).
Conclusion: Summative metrics-based feedback improved performance and lead to a safer and faster performance compared with standard instructions and seems a valuable educational tool in the early acquisition of temporal bone skills.
]]>Methods: Prospective study gathering validity evidence according to Messick’s framework. Four experts developed the CI Surgery Assessment Tool (CISAT). A total of 35 true novices (medical students), trained novices (residents) and CI surgeons performed two CI-procedures each in the Visible Ear Simulator, which were rated by three blinded experts. Classical test theory and generalizability theory were used for reliability analysis.
Results: The CISAT significantly discriminated between the three groups (p < 0.001). The generalizability coefficient was 0.76 and most of the score variance (53.3%) was attributable to the participant and only 6.8% to the raters. When exploring a standard setting for CI surgery, the contrasting groups method suggested a pass/fail score of 36.0 points (out of 55), but since the trained novices performed above this, we propose using the mean CI surgeon performance score (45.3 points).
Conclusion: Validity evidence for simulation-based assessment of CI performance supports the CISAT. Together with the standard setting, the CISAT might be used to monitor progress in competency-based training of CI surgery and to determine when the trainee can advance to further training.
]]>Method: In June 2020, 11 databases, including PubMed, were searched from inception through May 31, 2020. Eligible studies included the use of G-theory to explore reliability in the context of assessment of medical and surgical technical skills. Descriptive information on study, assessment context, assessment protocol, participants being assessed, and G-analyses were extracted. Data were used to map G-theory and explore variance components analyses. A meta-analyses was conducted to synthesize the extracted data on the sources of variance and reliability.
Results: Forty-four studies were included; of these, 39 had sufficient data for meta-analysis. The total pool included 35,284 unique assessments of 31,496 unique performances of 4,154 participants. Person variance had a pooled effect of 44.2% (95% confidence interval [CI] [36.8%-51.5%]). Only assessment tool type (Objective Structured Assessment of Technical Skills-type vs task-based checklist-type) had a significant effect on person variance. The pooled reliability (G-coefficient) was .65 (95% CI [.59-.70]). Most studies included D-studies (39, 89%) and generally seemed to have higher ratios of performances to assessors to achieve a sufficiently reliable assessment.
Conclusions: G-theory is increasingly being used to examine reliability of technical skills assessment in medical education but more rigor in reporting is warranted. Contextual factors can potentially affect variance components and thereby reliability estimates and should be considered, especially in high-stakes assessment. Reliability analysis should be a best practice when developing assessment of technical skills.
]]>Purpose: At graduation from medical school, competency in otoscopy is often insufficient. Simulation-based training can be used to improve technical skills, but the suitability of the training model and assessment must be supported by validity evidence. The purpose of this study was to collect content validity evidence for a simulation-based test of handheld otoscopy skills.
Methods: First, a three-round Delphi study was conducted with a panel of nine clinical teachers in otorhinolaryngology (ORL) to determine the content requirements in our educational context. Next, the authenticity of relevant cases in a commercially available technology-enhanced simulator (Earsi, VR Magic, Germany) was evaluated by specialists in ORL. Finally, an integrated course was developed for the simulator based on these results.
Results: The Delphi study resulted in nine essential diagnoses of normal variations and pathologies that all junior doctors should be able to diagnose with a handheld otoscope. Twelve out of 15 tested simulator cases were correctly recognized by at least one ORL specialist. Fifteen cases from the simulator case library matched the essential diagnoses determined by the Delphi study and were integrated into the course.
Conclusion: Content validity evidence for a simulation-based test of handheld otoscopy skills was collected. This informed a simulation-based course that can be used for undergraduate training. The course needs to be further investigated in relation to other aspects of validity and for future self-directed training.
]]>Purpose: Reliable assessment of surgical skills is vital for competency-based medical training. Several factors influence not only the reliability of judgements but also the number of observations needed for making judgments of competency that are both consistent and reproducible. The aim of this study was to explore the role of various conditions-through the analysis of data from large-scale, simulation-based assessments of surgical technical skills-by examining the effects of those conditions on reliability using Generalizability theory.
Method: Assessment data from large-scale, simulation-based temporal bone surgical training research studies in 2012-2018 were pooled, yielding collectively 3,574 assessments of 1,723 performances. The authors conducted generalizability analyses using an unbalanced random-effects design, and they performed decision studies to explore the effect of the different variables on projections of reliability.
Results: Overall, five observations were needed to achieve a Generalizability coefficient > 0.8. Several variables modified the projections of reliability: increased learner experience necessitated more observations (5 for medical students, 7 for residents, and 8 for experienced surgeons); the more complex cadaveric dissection required fewer observations than virtual reality simulation (2 vs. 5 observations); and increased fidelity simulation graphics reduced the number of observations needed from 7 to 4. The training structure (either massed or distributed practice) and simulator-integrated tutoring had little effect on reliability. Finally, more observations were needed during initial training when the learning curve was steepest (6 observations) compared with the plateau phase (4 observations).
Conclusions: Reliability in surgical skills assessment seems less stable than it is often reported to be. Training context and conditions influence reliability. The findings from this study highlight that medical educators should exercise caution when using a specific simulation-based assessment in other contexts.
]]>METHODS: A panel of fellowship-trained content experts in mastoidectomy was surveyed in relation to the 16 items of the assessment tool to determine the skills needed for supervised and unsupervised surgery. We examined the consensus score to investigate the degree of agreement among respondents for each survey item as well as additional analyses to determine whether the reported skill level required for each survey item was significantly different for the supervised versus unsupervised level.
RESULTS: Ten panelists representing different US training programs responded. There was considerable consensus on cut-off scores for each item and trainee level between panelists, with moderate (0.62) to very high (0.95) consensus scores depending on assessment item. Further analyses demonstrated that the difference between supervised and unsupervised skill levels was significantly meaningful for all items. Finally, minimum-passing scores for each item was established.
CONCLUSION: We defined performance standards for the cross-institutional mastoidectomy assessment tool using the Angoff method. These cut-off scores that can be used to determine when trainees can progress from performance under supervision to performance without supervision. This can be used to guide training in a competency-based training curriculum.
]]>