Speed
Less than half the time
2:35 per case with Arkangel AI versus 5:44 with traditional search, a reduction of about 55%.
Research stories
We randomly assigned clinical-year medical students to solve cases with or without Arkangel AI, while blinded specialists scored every answer. Students with the assistant answered better, in less than half the time, with half as many searches.
Imagine a timed open-book exam. By chance, half the students receive an assistant that does not supply the answer but finds the exact source; judges then review every response without knowing who had help. We did this with 83 students from four Colombian universities. They solved four outpatient cases with sixteen open questions using either Arkangel AI or their usual search methods. Two external, blinded specialists per field scored every answer on six clinical-validity criteria.
Students with the assistant scored better on every criterion, used less than half the time, and made half as many searches. Overall acceptability was 2.86 out of 3.
Read the original articleThe story
Imagine a timed open-book exam. Every student sees the same cases and may search anywhere. The only difference is that chance gives half of them a research assistant that does not provide the answer but finds the source and exact reference. The other half searches as usual through Google, PubMed, books, guidelines, or a colleague. Judges then review every answer without knowing who had the assistant.
More students are learning with artificial intelligence, yet most evaluations use multiple-choice tests, which barely resemble reasoning through a real case. A student on rotation does not select one of four options: they search, read, interpret, and answer in their own words.
We invited fourth-year or more senior students from four Colombian universities—Antioquia, los Andes, El Bosque, and CES—and randomly assigned them to two groups. Group A used Arkangel AI, which searches scientific sources and shows references for each response. Group B used usual search methods without AI. There were 83 participants: 43 with the assistant and 40 without it.
Everyone took the same test: four outpatient cases written by specialists in orthopedics, psychiatry, pediatrics, and obstetrics and gynecology. Each case included four open questions on diagnosis, management, evidence, and general knowledge. Responses were online and timed.
Two specialists in each field, all external to Arkangel AI, scored the answers without knowing each writer’s group. They assessed accuracy, agreement with guidelines, bias, currency, patient safety, and clinical relevance.
The experiment in numbers
Assignment was randomized in Excel, and both researchers and evaluators were blinded to each participant’s identity and group: a double-blind design.
What we found
Group A outperformed Group B on every validity criterion, all p<0.001. Mean total validity was 2.73 out of 3 with Arkangel AI versus 2.55 with traditional search. The largest relative gain was answer accuracy: 12.76%.
Efficiency changed too. Median time per case was 2 minutes 35 seconds with Arkangel AI versus 5 minutes 44 seconds, about 55% less. Median searches fell from 6 to 3 per case, about 50% less. Both differences had p<0.001.
The advantage held across every specialty and question type. Overall acceptability was 2.86 out of 3; confidence in model truthfulness received the lowest score, 2.65—a healthy skepticism.
Speed
2:35 per case with Arkangel AI versus 5:44 with traditional search, a reduction of about 55%.
Effort
3 searches per case versus 6. Students reached the source in fewer steps.
Consistency
The advantage held by specialty and question type; management and gynecology showed the largest differences.
Acceptability
Overall acceptability was 2.86 out of 3; confidence in truthfulness was 2.65.
What it means—and what it does not
The study supports a specific claim: in simulated outpatient cases with random assignment and blinded judges, students using a clinical search assistant with traceable sources wrote answers rated as more valid, in less than half the time and with half as many searches. They wrote open answers rather than selecting multiple-choice responses.
It does not prove that AI works for everything or everyone. The sample was modest and came from a few Colombian universities; the cases were fictional, outpatient, and non-urgent. The results represent academic clinical reasoning, not time-critical or inpatient decisions. There was no no-search control group, and a three-point scale has limited sensitivity.
Arkangel AI funded the project, and four of five authors are affiliated with the company. Mitigations included random assignment, double blinding, external case writers and evaluators, independent methodological review, and shared data and scripts for reproducibility.
The message is not that an assistant makes anyone a better physician. The useful question is whether it improves how real students search, read, and respond to clinical cases. This study offers a way to measure that and points toward validation with practicing physicians.
Paper
This double-blind randomized study enrolled fourth-year or more senior medical students at four Colombian universities. Excel randomization assigned 43 students to Arkangel AI-assisted search and 40 to traditional search without AI. Each solved four fictional outpatient cases with four open questions. The design contemplated up to 1,328 responses across 332 case units. Documented exclusions covered a platform failure, prohibited ChatGPT use, one incomplete test, and one duplicate questionnaire.
Two blinded external specialists per field applied six three-point Likert items. Mean total validity was 2.73 (95% CI 2.71–2.76) in Group A versus 2.55 (95% CI 2.52–2.59) in Group B, with p<0.001 across all domains. Median time was 2:35 (IQR 2:01–3:46) versus 5:44 (IQR 2:58–7:47), while median searches fell from 6 to 3; both p<0.001.
Limitations include the modest sample, few centers, fictional outpatient cases, no no-search control, heterogeneous traditional sources, the three-point scale, acceptability measured only in the intervention group, and most authors’ affiliation with Arkangel AI.