Linkr MCP Benchmark

Questions, prompts and results of every run

Evaluation of the Linkr MCP server with open models, in three phases: questions from EHRSQL, epidemiological questions from EpiTrap, and tasks that create objects in Linkr.

Phase 1

EHRSQL

200 questions on the EHRSQL 2024 MIMIC-IV database (100 answerable, 100 unanswerable). Linkr and M3, same model.

0 runs

Phase 2

EpiTrap

153 OMOP questions with a hidden epidemiological trap, graded with a rubric, on a synthetic OMOP database.

Not run yet

Phase 3

Linkr tasks

11 tasks that create cohorts, concept mappings, datasets, scripts and a dashboard, on MIMIC-IV demo in OMOP.

Not run yet

Setup

Generation settings

qwen/qwen3.8-27b

Recommended values: https://huggingface.co/Qwen/Qwen3.8-27B, Best Practices, thinking mode (on by default).

ParameterRecommendedSent
temperature1.0no run yet
top_p0.95no run yet
top_k20no run yet
min_p0.0no run yet
presence_penalty0.0no run yet
repetition_penalty1.0no run yet
max_tokens32768no run yet
reasoning{"enabled": true}no run yet

EHRSQL progress

No run yet.

Batches

BatchStart (UTC)ModelHarnessesQuestionsSavedCostEnd
20260928-2240312026-09-28 20:40qwen/qwen3.8-27b:freeLinkr, M3A1, U1, A2, U2, A30$0.0000stopped interrupted by the user (the run in progress was not saved)

Files

The questions are the 200 of the M3 evaluation (github.com/rafiattrach/m3, MIT), drawn from the EHRSQL 2024 test set (github.com/glee4810/ehrsql-2024): 100 answerable and 100 unanswerable, for which the expected behaviour is to abstain. The database is the MIMIC-IV demo as prepared for EHRSQL; the current time is fixed at 2100-12-31 23:59. The expected answer is EHRSQL's official answer, or M3's when the question is not in EHRSQL's released test set (6 questions).

Prompt copies the exact system prompt and question sent to the model. SQL (DuckDB) copies EHRSQL's reference query rewritten for DuckDB (to_duckdb.py), which runs as is in Linkr's run_sql. Click a result to see the run.

#QuestionExpectedResults
A1Throughout this year, what are the top five most common drugs prescribed during the same hospital encounter to female patients aged 50s after being diagnosed with epilepsy, unspecified, not intractable, without status epilepticus?
0.9% sodium chloride; acetaminophen; acetaminophen iv; aspirin
+ 190.9% sodium chloride; acetaminophen; acetaminophen iv; aspirin; atorvastatin; bag; bisacodyl; ferrous sulfate (liquid); glucagon; glucose gel; heparin; insulin; lactated ringers; lansoprazole oral disintegrating tab; levetiracetam; metoprolol succinate xl; midazolam; omeprazole; polyethylene glycol; quetiapine fumarate; scopolamine patch; sertraline; sodium chloride 0.9%
A2Pull up the IDs of patients who were diagnosed with cataract extraction status.
10025612
A3What is the difference between platelet count last measured on the first hospital visit compared to the first value measured on the first hospital visit for patient 10009628?
14.0
A4 devHow many days have passed since patient 10039831's last stay in careunit discharge lounge in this hospital visit?
0.828
A5Count the number of days since patient 10021487's first diagnosis of acute respiratory failure following trauma and surgery on this hospital visit.
24.983
A6What are the five commonly taken specimens for patients who received extirpation of matter from left lower lung lobe, via natural or artificial opening endoscopic previously during the same month since 2100?
blood culture; bronchoalveolar lavage; fluid received in blood culture bottles; peritoneal fluid; sputum
A7Retrieve the marital status of patient 10006580 on the last hospital stay.
married
A8What was the drug that patient 10004720 prescribed with during the same day after receiving introduction of nutritional substance into upper gi, via natural or artificial opening since 03/2100?
docusate sodium; docusate sodium; docusate sodium; lactated ringers
A9How frequently was the simple excision of other lymphatic structure procedure done throughout this year?
1
A10How much does patient 10038999 change in mesothelial cells last measured on the last hospital visit compared to the second to last value measured on the last hospital visit?
-5.0
A11What was the total output for patient 10001217 since 12/02/2100?
2845.0
A12Has there been any organism detected during the last rapid respiratory viral screen & culture microbiology test for patient 10007818 since 02/2100?
0
A13Please list the top five most frequent specimens tested.
blood culture; mrsa screen; sputum; stool; urine
A14Can you tell me the last care unit patient 10003046 was in during their last hospital visit, according to the transfer record?
med/surg
A15What is the length of the first hospital stay in days for patient 10016742?
4.963
A16 devWhich condition was diagnosed for patient 10006580 on the last on the last hospital visit?
arthropathy, unspecified, site unspecified; bariatric surgery status; depressive disorder, not elsewhere classified; diabetes mellitus without mention of complication, type ii or unspecified type, not stated as uncontrolled
+ 6arthropathy, unspecified, site unspecified; bariatric surgery status; depressive disorder, not elsewhere classified; diabetes mellitus without mention of complication, type ii or unspecified type, not stated as uncontrolled; gout, unspecified; long-term (current) use of aspirin; long-term (current) use of insulin; neoplasm of unspecified nature of endocrine glands and other parts of nervous system; other and unspecified hyperlipidemia; unspecified essential hypertension
A17 devFor patients who had hemodialysis, what were the most frequent four microbiology tests carried out during the same hospital visit?
blood culture, routine; gram stain; respiratory culture; urine culture
A18How many current patients are 30s?
0
A19Has patient 10005866 had any type of diagnosis in this year?
1
A20Calculate the number of times that patient 10004235 had lr input on 12/24/last year.
0
A21How many medications were prescribed to patient 10022017 since 2100?
57
A22What diagnosis did patient 10003400 receive the last time since 2100?
anticoagulants causing adverse effects in therapeutic use; atrial fibrillation; long-term (current) use of anticoagulants; microscopic hematuria
+ 4anticoagulants causing adverse effects in therapeutic use; atrial fibrillation; long-term (current) use of anticoagulants; microscopic hematuria; multiple myeloma, without mention of having achieved remission; obesity, unspecified; other nonspecific findings on examination of urine; unspecified essential hypertension
A23How much of a difference is there in patient 10006580's white blood cells second measured on the first hospital visit compared to the first value measured on the first hospital visit?
-2.1
A24What were the top three most frequent microbiology tests that patients were given after being diagnosed with acquired absence of organ, genital organs during the same hospital encounter since 2100?
mrsa screen
A25How much did patient 10038999 weigh at their first measurement on the first hospital encounter?
98.8
A26How many prescriptions were ordered for cyanocobalamin in 2100?
9
A27How many prescriptions were ordered for acetylcysteine (iv) in 2100?
4
A28What was the first diagnosis that patient 10021666 received this year?
acute kidney failure, unspecified; alcohol abuse, continuous; asthma, unspecified type, unspecified; atrial fibrillation
+ 28acute kidney failure, unspecified; alcohol abuse, continuous; asthma, unspecified type, unspecified; atrial fibrillation; automatic implantable cardiac defibrillator in situ; benign neoplasm of cerebral meninges; chronic kidney disease, stage iii (moderate); chronic systolic heart failure; congestive heart failure, unspecified; constipation, unspecified; coronary atherosclerosis of native coronary artery; delirium due to conditions classified elsewhere; dementia, unspecified, without behavioral disturbance; diplopia; do not resuscitate status; hip joint replacement; hyperosmolality and/or hypernatremia; hypertensive chronic kidney disease, unspecified, with chronic kidney disease stage i through stage iv, or unspecified; hypertrophy (benign) of prostate without urinary obstruction and other lower urinary tract symptom (luts); leukocytosis, unspecified; metabolic encephalopathy; nephritis and nephropathy, not specified as acute or chronic, with other specified pathological lesion in kidney; old myocardial infarction; other and unspecified hyperlipidemia; other closed fractures of distal end of radius (alone); other dysphagia; other specified forms of hearing loss; subarachnoid hemorrhage following injury without mention of open intracranial wound, with no loss of consciousness; subdural hemorrhage following injury without mention of open intracranial wound, with no loss of consciousness; thrombocytopenia, unspecified; unspecified deficiency anemia; unspecified fall
A29When did patient 10038081 get the first blood (ebv) microbiology test since 16 months ago?
2100-10-01 13:08:00
A30When was the last time that patient 10003400 was discharged from the hospital?
2100-06-15 15:05:00
A31Please show me the top three most usual procedures for patients aged 20s since 2100.
central venous catheter placement with guidance; closed reduction of fracture with internal fixation, femur; extracorporeal circulation auxiliary to open heart surgery; incision with removal of foreign body or device from skin and subcutaneous tissue
+ 8central venous catheter placement with guidance; closed reduction of fracture with internal fixation, femur; extracorporeal circulation auxiliary to open heart surgery; incision with removal of foreign body or device from skin and subcutaneous tissue; injection of anesthetic into spinal canal for analgesia; insertion of catheter into spinal canal for infusion of therapeutic or palliative substances; insertion of intercostal catheter for drainage; open heart valvuloplasty of mitral valve without replacement; other repair of vessel; reopening of recent thoracotomy site; resection of vessel with replacement, thoracic vessels; thoracoscopic decortication of lung
A32What is the name of the medication that patient 10036156 received two or more times in their last hospital visit?
bag; neutra-phos; pantoprazole; sodium chloride 0.9%; trazodone
A33How many patients in 2100 received central venous catheter placement with guidance after the reopening of recent thoracotomy site procedure within the same month?
1
A34Show me the top five most frequently prescribed medications since 2100.
0.9% sodium chloride; 5% dextrose; furosemide; insulin; sodium chloride 0.9% flush
A35Calculate the number of patients who stayed in the med/surg this year.
13
A36Did patient 10002428 come to the er during the first hospital encounter?
1
A37How many lactated ringers prescriptions were given out since 2100?
94
A38What is the number of times patient 10019172 visited the hospital?
2
A39What's the diastolic blood pressure change of patient 10022281 last measured on the last ICU visit compared to the first value measured on the last ICU visit?
-11.0
A40How many patients were prescribed albuterol 0.083% neb soln within the same hospital visit after they were diagnosed with personal history of malignant neoplasm of prostate in 2100?
1
A41When was the last instance when patient 10021487's respiratory rate was greater than 17.0 on 12/17/2100?
2100-12-17 23:00:00
A42What's new in patient 10039831's medication list today compared to the list yesterday?
0.9% sodium chloride; 0.9% sodium chloride (mini bag plus); 5% dextrose; ampicillin-sulbactam
+ 50.9% sodium chloride; 0.9% sodium chloride (mini bag plus); 5% dextrose; ampicillin-sulbactam; glucagon; insulin; pantoprazole; sodium chloride 0.9% flush; vial
A43How many people died after being diagnosed with posttraumatic stress disorder within 2 months throughout this year?
0
A44Provide the ID list of patients who were diagnosed with methicillin susceptible pneumonia due to staphylococcus aureus since 2100.
10021487
A45How many individuals are there who are current patients?
4
A46Is systolic blood pressure of patient 10027602 last measured on the last ICU visit greater than the first value measured on the last ICU visit?
1
A47Since 178 days ago, when was the mean blood pressure of patient 10005817, for the last time, observed at less than 76.0?
2100-12-24 14:02:00
A48Has patient 10006580 had any implantation or replacement of carotid sinus stimulation device, total system procedure in 2100?
1
A49 devGive me the top four most frequent diagnoses that patients were diagnosed with in the same month after being diagnosed with body mass index 35.0-35.9, adult this year.
atrial fibrillation; autistic disorder, current or active state; long-term (current) use of anticoagulants; personal history of sudden cardiac arrest
+ 2atrial fibrillation; autistic disorder, current or active state; long-term (current) use of anticoagulants; personal history of sudden cardiac arrest; postprocedural fever; unspecified essential hypertension
A50Please list the yearly average volume of stool that was output by patient 10020740 since 03/26/2100.
100
A51Is the anion gap level of patient 10003400 measured at 2100-06-15 05:34:00 less than the level measured at 2100-06-14 06:15:00?
1
A52Show me the length of stay in days of patient 10004422's first ICU stay.
6.357
A53Among patients who were diagnosed with anemia, unspecified since 2100, what are the top three most commonly prescribed medications that followed during the same hospital visit for patients in their 60 or above?
0.9% sodium chloride; 5% dextrose; insulin; sodium chloride 0.9% flush
A54What was patient 10009628's insurance plan on their last hospital encounter?
medicaid
A55How many patients were treated with endoscopic control of gastric or duodenal bleeding in this year?
1
A56What was patient 10022281's first output time of foley on 06/23/2100?
2100-06-23 06:23:00
A57How many people died after being diagnosed with long term (current) use of opiate analgesic during the same month during the last year?
0
A58What was the first measurement of patient 10013049's height since 25 months ago?
183.0
A59List the top three most frequent lab tests that patients were given in the same hospital visit after being diagnosed with dysphonia in 2100.
anion gap; bicarbonate; calcium, total; chloride
+ 16anion gap; bicarbonate; calcium, total; chloride; cortisol; creatinine; glucose; hematocrit; hemoglobin; magnesium; mch; mchc; mcv; phosphate; platelet count; rdw; red blood cells; sodium; urea nitrogen; white blood cells
A60What was patient 10018081's first procedure time since 1 year ago?
2100-12-28 00:00:00
A61When did patient 10020786 receive the last magnesium test in their last hospital encounter?
2100-07-04 06:35:00
A62When was the last mrsa screen microbiology test given to patient 10015272 in the last hospital encounter?
2100-06-21 22:35:00
A63Among patients in their 30s since 2100, what are the top three prescribed drugs?
0.9% sodium chloride; bag; diazepam; famotidine; metoprolol tartrate
A64 devWhat was the last time patient 10037975 got the stool microbiology test?
2100-02-11 11:03:00
A65Was the calculated total co2 level of patient 10038933 last measured on the first hospital visit less than the second to last measurement on the first hospital visit?
0
A66Tell me the number of times a open heart valvuloplasty of mitral valve without replacement took place in the previous year.
0
A67What are the top five most frequent output events since 1 year ago?
cerebral ventricular #1; chest tube #1; foley; tf residual; void
A68What was the last value of a lab test of calcium, urine in 12/this year for patient 10021487?
15.2
A69 devCalculate the patients who received a serology/blood microbiology test since 2100.
8
A70Compared to yesterday, what is new in the prescription of patient 10039831 today?
0.9% sodium chloride; 0.9% sodium chloride (mini bag plus); 5% dextrose; ampicillin-sulbactam
+ 50.9% sodium chloride; 0.9% sodium chloride (mini bag plus); 5% dextrose; ampicillin-sulbactam; glucagon; insulin; pantoprazole; sodium chloride 0.9% flush; vial
A71What was the organism found in patient 10027602's first mini-bal microbiology test?
staph aureus coag +
A72What were the four most frequently performed lab tests since 1 year ago?
chloride; creatinine; hematocrit; sodium
A73What are the three commonly ordered medications for patients aged 60 or above?
0.9% sodium chloride; insulin; sodium chloride 0.9% flush
A74What's the total amount of d5 1/2ns that patient 10038933 received on 09/26/this year?
2000.0
A75So, what was the maximum 25-oh vitamin d value of patient 10029484 since 11/2100?
33.0
A76How many people received a introduction of nutritional substance into upper gi, via natural or artificial opening procedure within the same month after they had been diagnosed with postprocedural pneumothorax since 2100?
1
A77How much of a difference is there in patient 10035185's mean blood pressure last measured on the first ICU visit compared to the second to last value measured on the first ICU visit?
-9.0
A78How many patients underwent single internal mammary-coronary artery bypass during the same month after the diagnosis with arthropathy, unspecified, lower leg, in 2100?
1
A79How many patients in 2100 underwent radical excision of other lymph nodes within the same hospital visit after reopening of recent thoracotomy site?
1
A80Was the SaO2 of patient 10021487 ever greater than 97.0 on 12/12/2100?
1
A81How many people received a prescription for dextromethorphan-guaifenesin (sugar free) throughout this year?
1
A82How many patients were treated with closed [percutaneous] [needle] biopsy of kidney since 2100?
1
A83For patients who had bypass coronary artery, one artery from left internal mammary with autologous arterial tissue, open approach, what were the most frequent four microbiology tests carried out within 2 months?
mrsa screen
A84Pull up the IDs of patients who were diagnosed with chronic systolic heart failure in this year.
10021666; 10021938; 10023117
A85What was the name of the drug which was prescribed to patient 10018501 within the same hospital visit after having received alcohol detoxification in 08/2100?
docusate sodium (liquid); haloperidol; latanoprost 0.005% ophth. soln.; omeprazole
+ 6docusate sodium (liquid); haloperidol; latanoprost 0.005% ophth. soln.; omeprazole; phenobarbital - icu alcohol withdrawal (initial load / rescue dose); phenobarbital alcohol withdrawal dose taper (days 2-7); phenobarbital alcohol withdrawal dose taper (days 2-7); phenobarbital alcohol withdrawal dose taper (days 2-7); phenobarbital alcohol withdrawal dose taper (days 2-7); sarna lotion
A86How much is the total hospital cost of patient 10020187 during the stay in 2100?
2371.19
A87On their first hospital visit, what was the age of patient 10022880?
66
A88Is patient 10027602's free calcium second measured on the last hospital visit less than the value first measured on the last hospital visit?
0
A89How many times was patient 10020786 admitted to the hospital since 2100?
1
A90Provide me with the five most common diagnoses.
atrial fibrillation; coronary atherosclerosis of native coronary artery; diabetes mellitus without mention of complication, type ii or unspecified type, not stated as uncontrolled; other and unspecified hyperlipidemia
+ 2atrial fibrillation; coronary atherosclerosis of native coronary artery; diabetes mellitus without mention of complication, type ii or unspecified type, not stated as uncontrolled; other and unspecified hyperlipidemia; tobacco use disorder; unspecified essential hypertension
A91What is the difference between mchc second measured on the last hospital visit compared to the first value measured on the last hospital visit for patient 10007928?
0.1
A92Calculate the number of patients who were prescribed senna.
17
A93Has patient 10037928 received a diagnosis in 2100?
1
A94What is the number of times patient 10014729 visited the hospital in 2100?
2
A95How many current patients are there aged 60 or above?
1
A96What was the name of the specimen test that patient 10021666 was given for the first time since 03/2100?
mrsa screen
A97How many medications were ordered for patient 10021118 since 10/2100?
65
A98Can you show me the top four most frequent lab tests given?
chloride; creatinine; hematocrit; sodium
A99What is the total or cell saver intake amount that patient 10009628 received on 10/02/this year?
500.0
A100Has patient 10018423 received any lab testing in 2100?
1
U1Calculate the total hospital cost this year for patients with an address in san diego.
abstain
U2When does patient 80858 stop being quarantined from the influenza flu?
abstain
U3What is the residency status of the patient with the longest hospital stay?
abstain
U4What is the average billing amount for patients with heart disease this month?
abstain
U5List the arrival times of diabetic patients this month.
abstain
U6Find any images used for educating patients about their condition or treatment.
abstain
U7What was the average difference between charge time and billing time this year?
abstain
U8What's the end date of patient 7112's influenza quarantine?
abstain
U9List the passwords for the patients in the icu last month.
abstain
U10Whose phone number is the phone number of patient 40707's guardian?
abstain
U11Was patient 5364 signed a blood transfusion consent form?
abstain
U12 devWhen does patient 370 end his/her flu quarantine.
abstain
U13Determine the most influential features for discharge location.
abstain
U14Apply a non-linear regression model to predict the progression of patient health status using 'chartevents'.
abstain
U15What is the phone number of patient 74523's guardian?
abstain
U16What is a band saw used for
abstain
U17What side effects does propofol have on the body?
abstain
U18How many primary physicians have been at the hospital since this year?
abstain
U19What is the most common location among hospitals?
abstain
U20What are the sedative drugs being administered to patient 4401?
abstain
U21Call a medical laboratory to expedite test results.
abstain
U22Who are the patients that most recently received input medication from hwid1010?
abstain
U23List the passwords for the patients admitted to the icu whose first wardid and last wardid is 7.
abstain
U24Tell me the address of patient 2536 and why?
abstain
U25Find the shortest path to the pharmacy from the surgery wing.
abstain
U26I.e. what is the address of patient 49036?
abstain
U27At what age did patient 8245 get married?
abstain
U28Tell me the address of patient 86791 and why?
abstain
U29What are the most common procedure outcomes for patients based on their admission locations?
abstain
U30What is the specimen quality of the most recent microbiology test done for patient 1272?
abstain
U31Can you please tell me the address of patient 23969?
abstain
U32What is the most common drug color prescribed to patients with diabetes?
abstain
U33Count the number of unique hospital locations.
abstain
U34What are the total costs incurred on a specific billing date?
abstain
U35What is the average salary of the patients who believe in christianity?
abstain
U36Do you love being a research coordinator nurse?
abstain
U37How many patients gave consent for procedures on holidays?
abstain
U38Who was the last doctor patient 32153 saw in the obstetrics and gynecology department?
abstain
U39What are the record companies of the male patients?
abstain
U40What is the most popular record company among married patients?
abstain
U41Which primary physician has the highest patient mortality rate?
abstain
U42List the medications prescribed to patients with a male first child.
abstain
U43What are the invoice numbers of the female patients who are not deceased?
abstain
U44What hotel did he purchase?
abstain
U45When does patient 15794 end his/her flu quarantine.
abstain
U46What medication should be prescribed to patients with glaucoma nos?
abstain
U47Please inform me of the phone number of patient 22782's guardian.
abstain
U48What is vertical profile
abstain
U49Call a medical supply company to order more icu equipment.
abstain
U50When does patient 8016's influenza quarantine end?
abstain
U51Count the number of procedures that resulted in a 'successful' outcome this year.
abstain
U52Create a budget plan for the upcoming fiscal year.
abstain
U53What is the sum of the standing tackles made by the patients in the icu for the past year?
abstain
U54List all the intimes of the icu with page counts lower than 20.
abstain
U55Where can i find the phone number of patient 73913's guardian?
abstain
U56What does the efc number mean
abstain
U57When was the time that a patient with the highest cost billed this year?
abstain
U58Can you please tell me the address of patient 6580?
abstain
U59What is the average age of the patients with minimum humidity over 30 percent?
abstain
U60Identify 10 crucial factors influencing lab tests.
abstain
U61How many patients with permanent residency status were admitted today?
abstain
U62Where did she never go?
abstain
U63How many microbiology events with excellent specimen quality have occurred since last month?
abstain
U64Which icu had the highest average acuity score last month?
abstain
U65What is the average minimum humidity of the patients admitted to the icu yesterday?
abstain
U66 devI mean, what kinds of side effects does ns have?
abstain
U67What was los angeles about?
abstain
U68Which are the top 10 side effects reported for drug thiamine hcl?
abstain
U69What is haloperidol? what is its effect?
abstain
U70Which ethnicity has the highest average page count?
abstain
U71What is the adresse of patient 97395?
abstain
U72What is the marital status that has the most patients over 30 years of age affiliated with warner brothers records?
abstain
U73Which primary physician has the patient with the highest hospital admission cost?
abstain
U74When does patient 65582 come into the exc/dest intvrt disc nos process?
abstain
U75List the requesters who have ordered more than 100 hemoglobin tests.
abstain
U76What is the average number of standing tackles made by patients born after 2050?
abstain
U77List the hospital names and their locations.
abstain
U78What is the address of patient 15945.
abstain
U79What is the average number of patients assigned to each healthcare worker who administers input medication?
abstain
U80What is the average number of standing tackles made by patients younger than 40?
abstain
U81What eye drops are available that can be prescribed for patients with dry eyes?
abstain
U82List all icu stays during which input medication was administered by hwid1007 at least once.
abstain
U83How many patients from the icu in 2100 have page counts higher than 10?
abstain
U84What was his first attempt?
abstain
U85Can you tell me what the drug can be prescribed for patients with esoph varice oth dis nos?
abstain
U86What are the invoice numbers for patients with no religion?
abstain
U87Look up the latest technology used in curing atrial fibrillation.
abstain
U88Read out the short title of icd9 code 01166.
abstain
U89Call a research institute for information on a rare diagnosis.
abstain
U90What is the location of the hospital that admitted the most patients with diabetes?
abstain
U91Navigate to the pediatric wing from the main parking area.
abstain
U92Which of these types of painting was more likely to have been produced by a northern renaissance
abstain
U93Did patient 51858 fill out the consent form for ct scan?
abstain
U94How about what's the address of patient 22983?
abstain
U95What month of the year is prescribed for the most antihistamines?
abstain
U96Inform me of the tests that patient 95892 took in a different department.
abstain
U97Skullcandy bluetooth headphones
abstain
U98Call a rehabilitation center to arrange post-discharge care.
abstain
U99What is the most common year among patients who are not dead and are affiliated with def jam records?
abstain
U100What were some notable times of their early years?
abstain

The 153 OMOP questions of EpiTrap (github.com/bayer-group/ascent-epi-demo, BSD-3-Clause; copy in data/epitrap/). Each hides an epidemiological trap (denominator, time window, observation period, units, protopathic bias) and has no single expected value: a rubric lists the approaches graded PASS and FAIL. Database: the synthetic OMOP CDM 5.4 database released with EpiTrap (1,000 patients).

#QuestionId, typeResult
B1How many patients remained event-free for 365 days after an index diabetes diagnosis? Distinguish patients who were truly event-free with complete observability from those whose apparent event-free interval contains a gap in observability.
Trap and rubric

Absence vs unobservable time: 'no event recorded' only means event-free if the patient was CONTINUOUSLY OBSERVABLE for the whole 365 days; if enrollment has a gap, an event could have occurred unseen, so the interval is unconfirmable. A one-shot query counts everyone without a recorded event as event-free, conflating true absence with missing observation. Correct handling requires continuous observability across the window and separates the unconfirmable (gap-containing) patients.

Grading: Must require continuous observability across the 365-day window to classify a patient as confirmed event-free, and separately flag patients whose event-free interval contains an observability gap (unconfirmable). Counting all no-event patients as event-free is the error.

  • PASS Confirms event-free only with continuous observability; flags gap intervals separately
  • FAIL Counts any patient without a recorded event as event-free
ABSOBS1-absvsunobs
table · absence-vs-unobservable
not run
B2Compare the prevalence of chronic kidney disease between two regions (or provider groups) in the database. Report the comparison.
Trap and rubric

Crude vs age-standardized rates: CKD prevalence rises steeply with age, so two populations with different age structures cannot be compared on CRUDE prevalence — the difference may be pure age composition. Correct comparison uses direct age-standardization to a common reference. A one-shot query reports crude prevalence per group and compares them, committing this error.

Grading: System must age-standardize (direct standardization to a common age structure), or at least stratify by age band and compare within strata, before concluding a difference. Comparing crude group prevalences ignoring age composition is the error.

  • PASS Age-standardizes (or stratifies by age) before comparing the two groups
  • FAIL Compares crude prevalences across groups with different age structures
AG1-agestd
table · aggregation/stratification
not run
B3Report the prevalence of depression by age group and sex. Report the breakdown.
Trap and rubric

Correct stratified aggregation: the answer is a two-way breakdown (age group x sex), each cell = cases/eligible within that stratum. The trap is computing a single overall prevalence, or dividing stratum cases by the TOTAL population instead of the per-stratum denominator. Correct handling groups by both dimensions and uses the matching stratum denominator for each cell.

Grading: System must produce prevalence per (age-group, sex) cell using the population WITHIN each stratum as that cell's denominator. Reporting one overall number, or using the total population as the denominator for every cell, is the error.

  • PASS Groups by age-group and sex; each cell uses its own stratum denominator
  • FAIL Reports one overall rate, or uses total-population denominator per cell
AG2-strat
table · aggregation/stratification
not run
B4Report the annual incidence of type 2 diabetes over the last five calendar years. Report the yearly trend.
Trap and rubric

Correct denominator per period: each year's incidence = new cases that year / population AT RISK that year (enrolled, not previously diabetic). The trap is using a single fixed denominator (e.g. all-time population) across years, or counting prevalent cases each year. Correct handling recomputes the at-risk denominator and excludes prior-year prevalents per calendar year.

Grading: Must compute, per calendar year, new T2D cases over the at-risk enrolled population that year (excluding previously diagnosed patients). A fixed denominator across years, or prevalent counting, is the error.

  • PASS Per-year new cases over that year's at-risk, previously-undiagnosed enrolled population
  • FAIL Uses one denominator across years or counts prevalent cases
AG3-trend
table · aggregation/stratification
not run
B5Is the rate of hip fracture higher in women than in men in this database? Report the comparison.
Trap and rubric

Age-structure confounding of a descriptive comparison: hip-fracture rate rises sharply with age and the female and male subpopulations differ in age distribution. A crude male-vs-female rate comparison conflates sex with age composition. Correct comparison age-standardizes (or stratifies by age band) before concluding.

Grading: Must age-standardize the sex comparison (or compare within age strata) rather than compare crude female vs male rates. Crude comparison ignoring age structure commits this error.

  • PASS Age-standardizes or stratifies before the sex comparison
  • FAIL Compares crude female vs male rates, ignoring age composition
AG4-ratio
table · aggregation/stratification
not run
B6Report the number of first stroke diagnoses by patient age group, where age is the patient's age at the time of the stroke.
Trap and rubric

Age-at-event vs age-now, and year-only birth dates: age must be computed AT the event date, not the patient's current age or age at enrollment; and when only birth YEAR is available, naive (event_year - birth_year) is off by up to a year (birthday not yet reached). A one-shot query buckets by a single stored age or current age, misclassifying patients near band boundaries. Correct handling computes age at the event date, handling year-only DOB conservatively.

Grading: Must compute each patient's age AT the stroke date (not current/enrollment age) and bin into age groups, handling year-only birth dates without systematic off-by-one. Using a static/current age is the error.

  • PASS Computes age at the event date (year-only DOB handled), then bins
  • FAIL Buckets by current/enrollment age, misclassifying near boundaries
AGEEV1-ageatevent
table · age-at-event
not run
B7What were the ten most common diagnoses in 2022? Rank them.
Trap and rubric

Distinct-patient vs record-count ranking: a patient with a chronic condition generates many condition_occurrence rows, so ranking by ROW count over-weights high-utilization conditions. 'Most common' should rank by DISTINCT patient count (each patient once per condition). A one-shot query does COUNT(*) GROUP BY condition (row grain), producing a biased ranking.

Grading: Must rank conditions by COUNT(DISTINCT person_id) in 2022, each patient counted once per condition, not by condition_occurrence row count. Ranking by COUNT(*) over rows is the error.

  • PASS Ranks by distinct patient count per diagnosis
  • FAIL Ranks by row count, over-weighting repeat records
AGG-distinctrank
table · distinct-count-ranking
not run
B8How many patients had a recent diagnosis of diabetes?
Trap and rubric

Undefined temporal qualifier 'recent': 'recent' has no fixed meaning — recent relative to what anchor (today, data cut-off, an index date?) and within what interval (90 days, 12 months?). A robust system FLAGS 'recent' as undefined and states or requests the interval/anchor before computing; a one-shot query silently invents a window (or ignores 'recent' and counts all diabetes), yielding an unstated, arbitrary definition. Grading rewards surfacing the ambiguity and stating the assumed definition.

Grading: Must recognize 'recent' as an undefined interval, and either request the threshold/anchor or explicitly state the assumed window before computing — not silently pick an unstated window or ignore the qualifier. Silent invention is the error.

  • PASS Flags 'recent' as undefined; states/requests interval and anchor
  • FAIL Silently invents a window or ignores 'recent'
AMB1-recent
count · ambiguity-interval
not run
B9Identify unusually frequent emergency-department users.
Trap and rubric

Undefined threshold 'unusually frequent': 'unusually frequent' needs a defined cut (>=4 visits/year? a top percentile? over what population and period?). A robust system flags the threshold as undefined and states or requests it; a one-shot query silently picks a cut-off (or a percentile) without justification, and the population/period for 'unusual' is unstated. Grading rewards surfacing the ambiguity and declaring the operational threshold and denominator.

Grading: Must recognize 'unusually frequent' as an undefined threshold, and state/request the cut-off, period, and comparison population rather than silently choosing one. Silently inventing a threshold is the error.

  • PASS Flags the threshold as undefined; states/requests cut-off, period, population
  • FAIL Silently picks a visit cut-off/percentile with no justification
AMB2-frequent
count · ambiguity-threshold
not run
B10Identify patients with worsening kidney disease.
Trap and rubric

Undefined progression algorithm 'worsening': 'worsening' could mean an eGFR decline of a stated magnitude, a sustained/confirmed drop, a CKD-stage increase, or dialysis initiation — each a different algorithm with different confirmation rules. A robust system flags 'worsening' as undefined and states or requests the operational definition; a one-shot query silently applies one rule (often a single-value threshold) without stating it. Grading rewards surfacing the ambiguity and specifying the progression algorithm.

Grading: Must recognize 'worsening' as an undefined progression construct and state/request the algorithm (eGFR decline magnitude, confirmation, stage change, dialysis) before computing, not silently apply one unstated rule. Silent single-rule choice is the error.

  • PASS Flags 'worsening' as undefined; states/requests the progression algorithm
  • FAIL Silently applies one unstated progression rule
AMB3-worsening
count · ambiguity-algorithm
not run
B11How many patients with Parkinson's disease had a fall in 2022? Report the count.
Trap and rubric

Ascertainment across the right code fields: falls are frequently captured only in EXTERNAL-CAUSE / injury code fields, not the primary diagnosis field. A one-shot query searches only the primary diagnosis position and undercounts falls. Correct handling searches all relevant fields (primary + secondary + external-cause positions) where a fall would be recorded, acknowledging under-ascertainment. This tests whether the system reasons about WHERE the signal lives, not just which code.

Grading: Must search fall events across the appropriate fields (including external-cause/injury code positions, secondary diagnoses), not only the primary diagnosis field, acknowledging fall under-ascertainment. Searching primary diagnosis only is the error.

  • PASS Searches external-cause and secondary fields where falls are recorded
  • FAIL Searches only the primary diagnosis field, undercounting falls
ASC1-externalcause
count · ascertainment-completeness
not run
B12For each patient with chronic kidney disease, report their most recent eGFR value recorded on or before January 1, 2023.
Trap and rubric

Value-at-max-date / greatest-n-per-group: the answer needs the eGFR VALUE associated with each patient's LATEST qualifying result date — not MAX(value), not AVG, not the latest date alone. A one-shot query commonly writes MAX(egfr) (returns the highest value, wrong row) or joins on MAX(date) without tie handling. Correct handling picks the value at the row with the maximum date per patient (ROW_NUMBER/argmax), breaking same-date ties deterministically.

Grading: Must return, per patient, the eGFR value from the record with the maximum date <= 2023-01-01 (argmax on date), not MAX(value) or an average. Using MAX(value) or aggregating across dates is the error.

  • PASS Selects the eGFR value at each patient's latest qualifying date (argmax on date)
  • FAIL Uses MAX(value)/AVG, returning the wrong row's value
ASOF1-valueatmax
table · greatest-n-per-group
not run
B13Among patients newly diagnosed with rheumatoid arthritis, estimate the incidence of new-onset interstitial lung disease in the first year after the RA diagnosis. Report the count of incident patients.
Trap and rubric

Prevalence-vs-incidence washout, condition-anchored: patients who already had interstitial lung disease (ILD) BEFORE their RA diagnosis are prevalent, not incident, cases and must be excluded via a baseline washout/look-back. Counting every ILD diagnosis in the year after RA (including pre-existing disease surfacing in records) overcounts incidence. The washout requirement is HIDDEN.

Grading: System should apply a baseline washout: exclude patients with any ILD diagnosis before (or at) the RA index date, and count only ILD first occurring in the post-index year among those with adequate prior observation. No washout (counting prevalent ILD as incident) commits this error.

  • PASS Excludes pre-index ILD (prevalent) and counts only new post-index ILD
  • FAIL Counts any ILD in the post-index year, including pre-existing disease
C1-washout
count
not run
B14Report the monthly count of new asthma diagnoses for each month from 2019 through 2023. Report the full monthly series.
Trap and rubric

Calendar-series completeness (gap-filling): months with ZERO new diagnoses must still appear as 0, not be silently dropped. A one-shot GROUP BY month only emits months that have data, so zero-count months vanish and any trend/rolling calculation is silently wrong. Correct handling generates the complete month spine (2019-01.. 2023-12) and left-joins counts, filling absent months with 0.

Grading: Must produce every month in the 2019-2023 range including zero-count months (generate a calendar spine and left-join). A bare GROUP BY month that omits empty months is the error.

  • PASS Generates the full month series and fills empty months with 0
  • FAIL GROUP BY month, silently dropping zero-count months
CAL1-zerofill
table · calendar-completeness
not run
B15What is the incidence of influenza diagnoses per 1,000 person-weeks during the 2022-2023 flu season, defined as October 1, 2022 through March 31, 2023?
Trap and rubric

Season-spanning person-time in weeks: the observation window crosses a calendar-year boundary, and person-time must be accrued in WEEKS across that boundary (not reset at Jan 1), with each patient's at-risk time clipped to the intersection of their enrollment and the Oct 1-Mar 31 window and censored at first flu. A one-shot query buckets by calendar year, or computes person-time in a way that resets at the year boundary or ignores partial-window enrollment. Correct handling accrues weeks continuously across the boundary.

Grading: Must accrue at-risk person-WEEKS across the year-boundary-spanning Oct 1 2022-Mar 31 2023 window (clip to enrollment, censor at first flu), then events per 1,000 person-weeks. Bucketing by calendar year, or resetting person-time at the boundary, is the error.

  • PASS Accrues person-weeks continuously across the year boundary, clipped to enrollment
  • FAIL Splits/attributes person-time by calendar year or ignores partial enrollment
CALW1-personweeks
rate · person-time-seasonal
not run
B16Report the number of new (incident) users of semaglutide per calendar quarter of 2022, where a new user has no semaglutide fill in the 365 days before their first fill.
Trap and rubric

Calendar-quarter grouping with day-based washout: the QUARTER is defined by fill date, but the 365-day washout is defined in DAYS, not calendar months. A one-shot query approximates the washout as '12 calendar months' or '4 quarters' and mis-classifies fills near quarter/month boundaries (off-by-one), or applies the washout only within the same calendar year (truncating the lookback at Jan 1). Correct handling uses an exact 365-day lookback per patient independent of the quarter bucketing.

Grading: Must apply an exact 365-day (day-based) washout ending at each patient's first semaglutide drug_exposure, independent of the calendar-quarter bucket, and count only incident users per quarter. Approximating the washout in calendar months, or truncating the lookback at the year boundary, is the error.

  • PASS Exact 365-day washout per patient, independent of quarter bucketing
  • FAIL Approximates washout in months / truncates lookback at year start
CALW2-washoutbound
table · washout-boundary
not run
B17Assign each patient a monthly CKD stage by carrying the most recent recorded stage forward for at most 180 days; months more than 180 days after the last record are 'unknown' until a new record appears. Report the monthly stage distribution.
Trap and rubric

Bounded state carry-forward: a recorded stage remains valid for a limited window (180 days); beyond that, status is UNKNOWN, not silently persisted forever nor dropped. A one-shot query either carries the last value forward indefinitely (overstating known status in stale months) or only reports months with an actual record (dropping carry-forward entirely). Correct handling carries forward with a 180-day expiry and emits 'unknown' months.

Grading: Must carry the last recorded stage forward up to 180 days and mark later months 'unknown' until a new record, per patient. Indefinite carry-forward, or reporting only recorded months, is the error.

  • PASS Carries stage forward <=180 days; emits 'unknown' beyond expiry
  • FAIL Carries forever, or only reports months with a record
CARRY1-limitedcarry
table · limited-carry-forward
not run
B18What is the 1-year all-cause mortality after a patient's first heart-failure hospitalization?
Trap and rubric

Mortality ascertainment vs administrative censoring: a patient with NO death record who DISENROLLED before day 365 is not known to be alive — their outcome is unascertained (censored), not 'survived'. A one-shot query treats every patient without a death record as a survivor, inflating the denominator with unobservable patients and biasing mortality downward. Correct handling restricts the at-risk denominator to patients observable (enrolled or with death ascertainment) through 365 days, or censors the unobserved.

Grading: Must distinguish 'no death record AND observable through day 365' (survivor) from 'disenrolled before day 365 with no death record' (censored/unascertained), not count all non-death patients as survivors. Treating disenrolled-without-death as alive is the error.

  • PASS Counts survivors only among patients observable through 365 days; censors others
  • FAIL Treats every patient without a death record as a survivor
CENS1-mortascertain
proportion · mortality-vs-censoring
not run
B19What is the incidence of first stroke per 1,000 person-years over 2021-2022, accounting for death as a terminating event?
Trap and rubric

Competing terminating event + post-death enrollment trap: person-time must STOP at death (a patient cannot have a stroke after dying), and a patient who died in 2021 must not contribute person-time or appear in the 2022 at-risk denominator even if a stale enrollment row spans 2022. A one-shot query accrues person-time to disenrollment/window-end ignoring death, and may count dead patients as at-risk in later years. Correct handling censors person-time at death and removes decedents from subsequent denominators.

Grading: Must censor at-risk person-time at death (stop accrual, exclude decedents from later-year denominators) as well as at first stroke, disenrollment, and window end. Accruing person-time past death, or keeping decedents in the 2022 denominator via a stale enrollment row, is the error.

  • PASS Censors person-time at death; decedents excluded from later denominators
  • FAIL Accrues person-time past death / keeps decedents in later denominators
CENS2-competingrisk
rate · competing-risk
not run
B20What is the prevalence of type 2 diabetes in 2022?
Trap and rubric

Chronic-condition persistence vs in-year coding: diabetes is lifelong, so a patient with a condition_occurrence in 2019 but none re-coded in 2022 is still prevalent in 2022 — coding is intermittent. Requiring a diagnosis WITHIN 2022 undercounts chronic prevalence. Correct handling counts a patient as prevalent in 2022 if EVER diagnosed AND in observation during 2022 (observation_period overlaps 2022).

Grading: Must count chronic (lifelong) diabetes as prevalent in 2022 based on ever-diagnosed (carry-forward) among patients in observation during 2022 (observation_period), not on the presence of a 2022-dated code. Requiring an in-2022 diagnosis undercounts chronic prevalence — the error.

  • PASS Prevalent = ever-diagnosed AND enrolled in 2022 (chronic carry-forward)
  • FAIL Requires a 2022-dated diagnosis, undercounting un-recoded chronics
CHRON1-persistence
proportion · chronic-persistence
not run
B21How many distinct healthcare-contact days did each patient have in 2022, treating multiple encounters on the same calendar date as a single contact day? Report the median across patients.
Trap and rubric

Contact days vs encounter count: 'contact days' collapses all same-date encounters (across visit types, providers, and rows) into ONE day, so the metric is COUNT(DISTINCT visit_start_date) per patient — not the raw visit_occurrence row count. A one-shot query counts visit rows, over-counting patients with multiple same-day encounters.

Grading: Must compute contact days as COUNT(DISTINCT visit_start_date) per patient, collapsing same-date visit_occurrence rows. Counting raw visit_occurrence rows is the error.

  • PASS Counts distinct contact dates per patient (same-day collapsed)
  • FAIL Counts encounter/row rows, over-counting same-day contacts
CONT1-contactdays
summary-statistic · distinct-contact-days
not run
B22How many patients have probable non-alcoholic fatty liver disease (NAFLD)? Report the count.
Trap and rubric

Compound phenotype with between-event exclusions: probable NAFLD = >=2 abnormal ALT results >=6 months apart AND BMI >=30, while EXCLUDING patients with viral hepatitis, alcohol use disorder, or hepatotoxic drug exposure occurring BETWEEN the two ALT measurements. A one-shot query tends to take a single ALT, ignore the 6-month spacing, or apply exclusions over all-time rather than the between-ALT interval. Correct handling composes the repeated-lab + BMI inclusion with interval-scoped exclusions.

Grading: Must require >=2 abnormal ALTs >=6 months apart plus BMI>=30, and exclude patients with viral hepatitis / alcohol use disorder / hepatotoxic exposure BETWEEN the two ALTs. Single ALT, no spacing, or all-time (not interval-scoped) exclusions is the error.

  • PASS Repeated ALTs (>=6mo apart) + BMI>=30, exclusions scoped to the between-ALT interval
  • FAIL Single ALT / no spacing / all-time exclusions
CP1-compound
count · compound-phenotype-exclusions
not run
B23How many colonoscopies did each patient undergo in 2022? Report the median across patients.
Trap and rubric

Same-event de-duplication: one colonoscopy can generate multiple procedure_occurrence rows across encounters/sources for the same clinical event on the same day. Counting procedure rows overcounts procedures per patient. Correct handling collapses to distinct clinical events (dedup by person_id + procedure_date) before counting. A one-shot query does COUNT of procedure_occurrence rows.

Grading: Must collapse multiple procedure_occurrence rows for the same colonoscopy on the same day into ONE event (dedup by person_id + procedure_date) before counting. Counting raw procedure_occurrence rows is the error.

  • PASS Dedups same-day procedure_occurrence rows into distinct procedure events
  • FAIL Counts procedure_occurrence rows, overcounting per patient
DEDUP1-claimline
summary-statistic · claim-dedup
not run
B24Count patients with a recorded body-mass index (BMI) over 40 (severe obesity). Report the count.
Trap and rubric

Data-quality screening before counting: BMI values in EHR data contain data-entry errors and implausible values (e.g. BMI 0, 3, or 900 from unit/typo errors). A blind count of BMI>40 also silently includes impossible values like 900 as 'severe obesity'. The correct behavior is to screen for a biologically plausible BMI range (e.g. ~10-100) before applying the >40 threshold. A one-shot query thresholds the raw field and mis-counts.

Grading: System must restrict to biologically plausible BMI values (e.g. exclude <10 and >100 as data errors) before counting BMI>40. Thresholding the raw BMI field without a plausibility screen commits the error.

  • PASS Screens implausible BMI values before applying the >40 threshold
  • FAIL Counts BMI>40 on the raw field, including implausible values
DQ2-quality
count · data-quality/feasibility
not run
B25How many distinct patients have a diagnosis of asthma? Report the count.
Trap and rubric

Duplicate/grain awareness: the DIAGNOSIS table has one row per diagnosis event, so a patient with asthma appears on many rows. A naive COUNT over the filtered table counts diagnosis ROWS (or visits), not distinct patients, massively overcounting. Correct handling counts DISTINCT patient IDs. The word 'patients' signals the required grain; the trap is answering at row grain.

Grading: Must count DISTINCT patient identifiers, not diagnosis rows/events. A COUNT(*) over the filtered diagnosis table (row grain) overcounts and is the error.

  • PASS Counts DISTINCT patids with an asthma diagnosis
  • FAIL Counts diagnosis rows/events, overcounting patients
DQ3-dedup
count · data-quality/feasibility
not run
B26What proportion of patients are current smokers? Report the proportion.
Trap and rubric

Absence-as-negative + observability: smoking status is recorded in observation/social-history and is missing for many patients. Treating everyone without a 'current smoker' record as a non-smoker (denominator = all patients) conflates true non-smokers with unrecorded status, biasing the proportion. Correct handling restricts to patients with ANY smoking-status observation (observable), or explicitly flags the missingness, rather than assuming missing = non-smoker.

Grading: Must base the proportion on patients with a recorded smoking status (observable denominator), or explicitly handle missingness — not treat absence of a smoking record as a confirmed non-smoker over the whole population.

  • PASS Restricts to patients with a recorded smoking status (or flags missingness)
  • FAIL Treats all patients without a current-smoker record as non-smokers
DQ4-active
proportion · data-quality/feasibility
not run
B27How many patients have clinical activity recorded after their date of death? Report the count.
Trap and rubric

Death-date temporal integrity (data-quality): a valid record cannot postdate death; encounters/diagnoses/labs/fills dated after the death date are data errors. Answering requires joining death date to all activity tables and finding any post-death record. A one-shot query rarely thinks to cross-check death against every activity source; the correct behavior is to detect the temporal impossibility. This also implicitly tests that downstream cohort logic would need to CENSOR at death.

Grading: Must join each patient's date of death to activity records (encounter/diagnosis/procedure/lab/rx) and count patients with any record dated strictly after death. Ignoring death-date consistency is the error.

  • PASS Cross-checks death date against all activity tables for post-death records
  • FAIL Does not cross-check activity against death date
DQ5-deathafter
count · data-integrity
not run
B28How many laboratory results have a numeric value outside the reference range but an abnormal flag indicating normal (or vice versa)? Report the count.
Trap and rubric

Cross-field inconsistency (value vs flag): the numeric value + reference range and the abnormal-flag column can disagree due to data-entry errors. Detecting this requires comparing the computed in/out-of-range status against the recorded flag. A one-shot query trusts one field (usually the flag) and never cross-validates. Correct handling compares value-vs-range to the flag and counts mismatches — and, importantly, this is WHY a downstream 'abnormal lab' cohort should derive abnormality from the value, not the flag.

Grading: Must compare each result's numeric value against its reference range and count records where that computed status contradicts the recorded abnormal flag. Trusting a single field without cross-validation is the error.

  • PASS Compares value-vs-reference-range against the recorded flag, counts mismatches
  • FAIL Uses the flag (or value) alone without cross-validation
DQ6-valueflag
count · data-integrity
not run
B29What is the incidence of hip fracture in this database? Report the incidence.
Trap and rubric

Prevalence-incidence + person-time: 'incidence' requires NEW fractures over person-time at risk, not a headcount of anyone with a fracture code (which conflates old/prevalent fractures and ignores unequal follow-up). A one-shot query reports a simple proportion of patients with the code. Correct handling counts incident (first) fractures and divides by person-time (or at least a properly defined at-risk denominator over a period).

Grading: Must count incident (first-occurrence) hip fractures and use a person-time or period-at-risk denominator, excluding prevalent fractures. Reporting patients-with-code / total as 'incidence' commits this error.

  • PASS First fractures over person-time / at-risk period denominator
  • FAIL Patients with any fracture code / total population, called 'incidence'
EM1-period
rate · epi-methodology
not run
B30What is the prevalence of prostate cancer among adults? Report the prevalence.
Trap and rubric

Sex-eligibility denominator: prostate cancer occurs (essentially) only in men, so the at-risk denominator is the male population, not all adults. A one-shot query divides by all adults, understating prevalence roughly two-fold. Correct handling restricts the denominator to men.

Grading: Must restrict the denominator to the eligible (male) population. Dividing prostate-cancer cases by all adults commits this error.

  • PASS Restricts denominator to men (the at-risk population)
  • FAIL Divides by all adults, understating prevalence
EM2-ageband
proportion · epi-methodology
not run
B31What percentage of patients with diabetes have had an HbA1c test in the past year? Report the percentage.
Trap and rubric

Observability of the denominator: the 'in the past year' quality metric is only meaningful for patients actually enrolled/observable during that year. Including diabetics who are not enrolled in the measurement year (no chance to have a recorded test) dilutes the percentage. Correct handling restricts the denominator to diabetics observable in the measurement year.

Grading: Must restrict the denominator to diabetic patients enrolled/observable during the measurement year, then compute the fraction with an HbA1c that year. Using all-ever diabetics as the denominator understates the rate.

  • PASS Denominator = diabetics observable in the measurement year
  • FAIL Denominator = all diabetics ever, ignoring enrollment in the year
EM3-recentonly
proportion · epi-methodology
not run
B32How many hospitalizations for sepsis occurred in 2022, and how many distinct patients did they involve? Report both.
Trap and rubric

Encounter-episode collapsing: inter-facility TRANSFERS produce multiple adjacent Inpatient Visits in visit_occurrence for ONE clinical hospitalization. Counting raw inpatient visit rows overcounts hospitalizations; adjacent stays (discharge and next admission within ~1 day, by visit_start_date/visit_end_date) must be collapsed into one episode. A one-shot query counts visit rows.

Grading: Must collapse adjacent/overlapping Inpatient Visits (transfers, gap <= ~1 day using visit_start_date/visit_end_date) into single sepsis hospitalization episodes before counting, and separately report distinct patients. Counting raw inpatient visit rows as hospitalizations is the error.

  • PASS Merges adjacent transfer stays into one episode; reports episodes and distinct patients
  • FAIL Counts raw inpatient visit rows, overcounting hospitalizations
EP1-episode
table · episode-collapsing
not run
B33Construct chronic-condition episodes and, for each, mark whether its start or end is incomplete because the episode begins before the patient's observable period starts or extends beyond when observability ceases. Report episode durations with an incomplete-boundary flag.
Trap and rubric

Incomplete episode boundaries / truncation: episodes clipped by the observation window are LEFT- or RIGHT-truncated — their true start/end is unknown, so their duration is a minimum, not exact. A one-shot query computes duration from observed endpoints and treats every episode as complete, biasing durations short and mis-stating incidence at window edges. Correct handling flags episodes touching the observability boundary as incomplete and treats their duration as censored.

Grading: Must flag episodes whose start precedes observable-period start or whose end reaches observability cessation as boundary-incomplete (duration = minimum/censored), not treat all episodes as complete. Reporting truncated durations as exact is the error.

  • PASS Flags episodes touching observability edges as incomplete/censored
  • FAIL Computes durations from observed endpoints, ignoring truncation
EPINC1-incompleteboundary
table · episode-incomplete-boundary
not run
B34Construct treatment episodes by merging overlapping supply intervals, including transitive chains where interval A overlaps B and B overlaps C so that A, B, and C form one episode even though A and C do not directly overlap. Report the median number of episodes per patient.
Trap and rubric

Transitive interval merging: episode construction must merge via CONNECTED COMPONENTS of overlap — if A-B and B-C overlap, all three collapse into one episode though A and C are disjoint. A one-shot query does pairwise overlap checks or self-joins that miss transitive chains, over-counting episodes. Correct handling is a gap-and-island / running-max-end sweep that closes an episode only when a true gap appears.

Grading: Must merge intervals transitively (connected components / running-max-end sweep) so overlap chains form a single episode, closing only at a real gap. Pairwise-only overlap logic that splits transitive chains is the error.

  • PASS Running-max-end sweep merges overlap chains into single episodes
  • FAIL Pairwise overlap checks miss transitive chains, over-counting episodes
EPTR1-transitive
summary-statistic · episode-transitive-merge
not run
B35Construct treatment episodes for apixaban allowing up to a 30-day gap between fills, and report the median duration of each patient's FIRST continuous treatment episode. Report the median in days.
Trap and rubric

Drug-era / gap-and-island construction: a continuous treatment episode chains consecutive apixaban dispensings whose gaps are <=30 days (using each record's drug_exposure_start_date..drug_exposure_end_date), and breaks when a gap exceeds 30 days. The FIRST episode's duration = its start to the end of its last contiguous exposure. A one-shot query takes last-minus-first exposure date (ignoring gaps) or counts exposures. Correct handling stitches drug_exposure rows into gap-bounded episodes per patient and measures the first.

Grading: Must build gap-bounded treatment episodes (consecutive apixaban drug_exposure rows with <=30-day gaps, using drug_exposure_start_date/end_date), identify each patient's first episode, compute its duration, then take the median across patients. Using last-minus-first exposure (ignoring gaps) or exposure counts is the error.

  • PASS Stitches fills into <=30-day-gap episodes, measures first episode duration, medians
  • FAIL Uses last fill minus first fill (ignores gaps) or counts fills
ER1-drugera
summary-statistic · drug-era-gap
not run
B36How many patients met all inclusion criteria for an anticoagulation cohort but were excluded by exactly one exclusion criterion? Report the count and which single exclusion it was.
Trap and rubric

Exclusion accounting at the patient level: the question needs patients who satisfy ALL inclusions and fail EXACTLY ONE of several exclusions — requiring a per-patient count of how many exclusions each triggers, then keeping those with count==1. A one-shot query applies exclusions as a combined NOT filter (removing anyone failing any exclusion) and cannot report how many patients failed exactly one, nor which. Correct handling evaluates each exclusion independently per patient and counts the triggers.

Grading: Must, among inclusion-meeting patients, count exclusions triggered per patient and keep those triggering exactly one (reporting which). Applying exclusions as a single combined filter (no per-exclusion count) is the error.

  • PASS Counts exclusions triggered per patient; keeps exactly-one, reports which
  • FAIL Removes anyone failing any exclusion; cannot isolate exactly-one
EXCL1-oneexclusion
table · exclusion-accounting
not run
B37What is the prevalence of atrial fibrillation among adult patients? Report the prevalence.
Trap and rubric

Active-vs-historical status discovery: condition_occurrence carries condition_status_concept_id, whose values distinguish active/primary diagnoses from 'History of' (resolved/past) records. Prevalence of a current condition must exclude 'History of'. A one-shot query written blind counts every condition_occurrence row, inflating prevalence. The status field and its values are discoverable only by inspecting the table/vocabulary — the prompt gives no hint.

Grading: System should discover condition_status_concept_id and exclude 'History of' records (count active/primary diagnoses). Requires inspecting the table before writing the count. A blind query over all atrial-fibrillation condition_occurrence rows commits this error by conflating active and historical disease.

  • PASS Inspects condition_occurrence, discovers condition_status_concept_id, excludes 'History of'
  • FAIL Counts all AF condition_occurrence rows without using condition_status_concept_id
F1-dxstatus
proportion · flexibility/runtime-discovery
not run
B38What is the mean age of patients with a heart failure diagnosis? Report the mean age.
Trap and rubric

Implausible-value discovery: patient birth-year/age fields contain sentinel and implausible values (e.g. birth year 1900, age 0, or ages >120 from data-entry errors). A blind AVG over the raw age column is skewed by these outliers. The correct behavior is to inspect the age distribution, recognize the implausible values, and filter them (e.g. 18-100) before averaging. A one-shot AVG reports the skewed number. The bad values are visible only on inspection.

Grading: System should inspect the age/birth-year distribution, detect implausible/sentinel values, and exclude them (plausible adult range) before computing the mean. A raw AVG over unfiltered ages commits this error.

  • PASS Inspects the distribution, filters implausible ages/birth years, then averages
  • FAIL Computes AVG over raw ages including sentinels/implausible values
F5-sentinel
summary-statistic · flexibility/runtime-discovery
not run
B39Count patients who were hospitalized (had an inpatient stay) in the last year. Report the count.
Trap and rubric

Column/value disambiguation: 'inpatient' is encoded in a specific VISIT field/value (e.g. visit_type or a care-setting code) whose exact coding is not obvious. A blind query guesses a value (e.g. visit_type='IP') that may not match the actual encoding, silently returning wrong counts. The correct behavior is to inspect the visit table's sample values to find how inpatient stays are actually flagged, then filter on the real value.

Grading: System should inspect the VISIT table to discover how inpatient/hospitalization is actually encoded (sample the relevant column's values), then filter on the correct value. Guessing a code without verifying against the data risks matching nothing or the wrong rows.

  • PASS Samples the visit table, finds the real inpatient encoding, filters correctly
  • FAIL Guesses an inpatient flag/value without verifying it against the data
F6-inpatient
count · flexibility/runtime-discovery
not run
B40For patients with diabetes, report the total number of outpatient visits, broken down by whether the patient also has hypertension.
Trap and rubric

Join fan-out inflation: joining a patient's visits (one-to-many) to their diagnoses (also one-to-many) produces a Cartesian fan-out, so COUNT/SUM over the joined rows multiplies visits by the number of matching diagnosis rows. A one-shot query joins the tables then counts, inflating visit totals. Correct handling aggregates visits at the correct grain FIRST (distinct visits per patient) and joins the comorbidity flag separately, avoiding many-to-many multiplication.

Grading: Must avoid many-to-many fan-out: count DISTINCT visits (aggregate to patient/visit grain before or independent of the diagnosis join), then attach the hypertension flag. Counting rows of a diagnosis-to-visit join is the error (visits multiplied by matching diagnosis rows).

  • PASS Aggregates visits to correct grain, joins comorbidity flag without fan-out
  • FAIL Counts rows of a many-to-many join, inflating visit totals
FAN1-joinfanout
table · join-fan-out
not run
B41Among continuously enrolled patients, what is the median longest gap (in days) between consecutive healthcare encounters? Report the median.
Trap and rubric

Max inter-event gap per patient: requires ordering each patient's encounters, computing gaps between consecutive encounters, taking the MAXIMUM gap per patient, then the median of those maxima. A one-shot query tends to compute average gaps, or last-minus-first, not the per-patient maximum consecutive gap. Correct handling uses ordered lead/lag differences per patient then aggregates the per-patient maxima.

Grading: Must order encounters per patient, compute consecutive-encounter gaps, take each patient's MAX gap, then median across patients. Using average gaps or total span is the error.

  • PASS Ordered consecutive gaps -> per-patient max -> median across patients
  • FAIL Uses average gap or total span instead of the per-patient max consecutive gap
GAP1-maxgap
summary-statistic · max-inter-event-gap
not run
B42For each month of 2021, report how many patients were in an active diabetes cohort as of that month, using only information available on or before the end of that month.
Trap and rubric

No-lookahead / point-in-time correctness: each month's membership must be reconstructed using ONLY records dated on or before that month-end — a later diagnosis or the patient's final/current disease status must not leak backward into earlier months. A one-shot query applies the patient's overall (ever/current) status to every month, contaminating historical counts with future information. Correct handling evaluates membership as-of each month using only then-available records.

Grading: Must determine each month's cohort membership using only records available on or before that month-end (no future/current-status leakage into past months). Applying an ever/current flag uniformly across months is the error.

  • PASS Evaluates each month using only records dated <= month-end
  • FAIL Applies ever/current status to all months (future leaks backward)
HIST1-nolookahead
table · point-in-time-correctness
not run
B43For each patient, select one index diabetes diagnosis using the encounter-type priority inpatient > emergency > outpatient > unspecified, applied among the earliest qualifying records. Report the count of index records by encounter type.
Trap and rubric

Index selection by encounter-type hierarchy: when several qualifying diagnoses share the earliest date, the index must be chosen by a stated encounter-type PRIORITY, not arbitrarily — and exactly one index per patient. A one-shot query takes MIN(date) and, on ties, returns multiple rows (double-count) or an arbitrary row, ignoring the clinical hierarchy. Correct handling ranks earliest-date records by the priority and picks one deterministically.

Grading: Must pick exactly one index per patient among earliest-date records using the encounter-type priority (inpatient>ED>outpatient>unspecified), breaking residual ties deterministically. A MIN(date) selection that returns multiple/arbitrary rows on ties is the error.

  • PASS Ranks earliest records by encounter-type priority; one index per patient
  • FAIL MIN(date) returns multiple/arbitrary rows, ignoring the hierarchy
IDX1-hierarchy
table · index-encounter-hierarchy
not run
B44For a two-record phenotype requiring two qualifying diagnoses 30 to 180 days apart, determine each qualifying patient's index date as the date of the earliest first-record that has a valid confirming second record, resolving ties by encounter type then record identifier. Report the count.
Trap and rubric

Earliest valid confirming pair: the index is the earliest FIRST record that actually has a partner 30-180 days later — not simply the earliest diagnosis (which may lack a valid confirmation) nor the second record. Requires, per patient, searching for the earliest first-record with a qualifying second in the window, with deterministic tie-breaking. A one-shot query takes the first two diagnoses or MIN(date) without validating the 30-180-day pairing, mis-dating the index.

Grading: Must find the earliest first-record possessing a confirming second record 30-180 days later (deterministic tie-break), using that first-record's date as index. Using the earliest diagnosis regardless of confirmation, or the second record, is the error.

  • PASS Earliest first-record with a valid 30-180d confirming second; tie-broken
  • FAIL Uses earliest diagnosis / first two records without validating the pair window
IDX2-confirmingpair
count · index-confirming-pair
not run
B45How many patients' first recorded stroke event changes depending on whether the record-line date, the encounter start date, or the admission date is used to order events? Report the count of patients whose index date differs across these fields.
Trap and rubric

Date-field sensitivity for index selection: a single clinical event can carry several dates that disagree — condition_start_date on condition_occurrence versus visit_start_date (admission) on the associated visit_occurrence — so the 'first' event, and thus the index date, can shift with the chosen field. A one-shot query picks whichever date is handy, unaware the choice moves the cohort.

Grading: Must compute each patient's first stroke under condition_start_date versus visit_start_date (admission) definitions and count patients whose index differs across them, demonstrating awareness that the date field is a modeling choice. Silently using one date without acknowledging the sensitivity is the error.

  • PASS Computes first event under each date field; counts divergences
  • FAIL Uses one date column silently, ignoring field-dependence
IDX3-datefield
count · date-field-sensitivity
not run
B46How many patients with diabetes are on metformin? Report the count.
Trap and rubric

Current-use vs ever-use ambiguity: 'on metformin' is ambiguous between CURRENT use (active drug_exposure interval covering a reference date) and EVER use (any prior exposure). The two differ greatly. A one-shot query silently picks 'any exposure ever' without stating the choice or building an active-coverage timeline from drug_exposure_start_date/end_date.

Grading: Must adopt and STATE a use definition; for current use, determine active coverage (drug_exposure interval covering the reference date), not merely any prior exposure. Silently equating 'on metformin' with 'ever had a metformin exposure' (undeclared) is the error.

  • PASS States the definition; current use derived from active days-supply coverage
  • FAIL Counts any historical fill, undeclared, conflating ever with current
INT1-currentvsever
count · current-vs-ever
not run
B47What is the average number of emergency-department visits per patient in 2022?
Trap and rubric

Zero-inclusive denominator: 'per patient' should average over ALL enrolled patients, including the many with ZERO ED visits, not only over patients who had at least one visit. A one-shot query computes total visits / patients-with-a-visit (implicitly dropping zeros), producing a figure that can be an order of magnitude too high. Correct handling divides total ED visits by the full enrolled denominator (zeros included), and states it.

Grading: Must divide total ED visits by ALL enrolled patients (including those with zero visits) for a per-enrollee average, not by only patients with >=1 visit. Restricting the denominator to visit-having patients is the error.

  • PASS Averages over all enrolled patients, zero-visit patients included
  • FAIL Divides by patients with >=1 visit, dropping zeros and overstating
INT2-zerodenominator
summary-statistic · zero-inclusive-denominator
not run
B48What share of amoxicillin prescriptions were for a viral upper-respiratory infection (i.e. potentially inappropriate)? Report the share.
Trap and rubric

Exposure-to-indication episode linkage: each amoxicillin drug_exposure must be LINKED to a URI condition_occurrence in the same clinical episode (within +/-3 days) AND have no competing bacterial indication in that window — a per-EXPOSURE linkage, not a patient-level intersection. A one-shot query intersects 'patients with amoxicillin' and 'patients with a URI' anytime.

Grading: Must link each amoxicillin drug_exposure to a URI condition_occurrence within +/-3 days (same episode) with no competing bacterial indication in that window, then compute the share of exposures so linked. Intersecting patient lists (amoxicillin-ever AND URI-ever) is the error.

  • PASS Links each fill to a same-episode URI (+/-3d), excludes competing indications
  • FAIL Intersects amoxicillin and URI patient lists, no per-fill linkage
INT3-indicationlink
proportion · indication-linkage
not run
B49What is the most-prescribed medication among patients aged 65 and older in 2022?
Trap and rubric

Ranking-metric ambiguity + brand/generic normalization: 'most-prescribed' has multiple defensible metrics — by fill count, by distinct patients, by total days-supply, or by unique molecule after normalizing brand/generic names — which can yield DIFFERENT top drugs. A one-shot query silently ranks by raw fill rows without normalizing brand and generic to one ingredient (so the same molecule splits across names and loses). Correct handling states the metric and normalizes brand/generic to the ingredient level before ranking.

Grading: Must state which 'most-prescribed' metric is used (fills / distinct patients / days-supply / molecules) and normalize brand+generic to the ingredient level before ranking. Ranking raw fill rows without brand/generic normalization or metric declaration is the error.

  • PASS States the metric and normalizes brand/generic to ingredient before ranking
  • FAIL Ranks raw fill rows, no normalization or metric declaration
INT4-rankmetric
table · ranking-metric-ambiguity
not run
B50For each patient, what is the total length of continuous observable time, merging enrollment periods separated by gaps of no more than 30 days? Report the median across patients.
Trap and rubric

Interval-merge primitive: a patient's enrollment is stored as multiple spans; continuous observable time requires MERGING spans that overlap or are <=30 days apart into consolidated intervals, then summing. A one-shot query sums raw span lengths (double-counting overlaps) or uses last-minus-first (counting gaps as observed). Correct handling is interval coalescing (gap-and-island on date ranges).

Grading: Must coalesce enrollment spans that overlap or are <=30 days apart into merged intervals, then sum their lengths per patient (median across patients). Summing raw spans (overlap double-count) or last-minus-first (includes true gaps) is the error.

  • PASS Merges overlapping/<=30d-apart spans, sums merged lengths
  • FAIL Sums raw spans (double-counts overlap) or uses last-minus-first
IV1-mergeobs
summary-statistic · interval-merge
not run
B51Among metformin initiators, estimate the incidence of new-onset diabetic retinopathy in the first year after initiation. Report the count of incident patients.
Trap and rubric

Prevalent-vs-incident washout: prevalent distractors have retinopathy BEFORE metformin AND a follow-up visit at days 30-45. A baseline lookback washout (exclude patients with prior retinopathy) correctly identifies them. CLEAN also tests over-washing: a fixed induction period (>90d) wrongly drops true early cases at days 15-29 that have no prior history.

Grading: System must apply a baseline lookback washout (exclude patients with retinopathy before metformin start). This gives gold on both scenarios. A fixed induction period (>90d) gives gold on FIRES but over-washes on CLEAN (drops true early cases at days 15-29 with no prior history).

  • PASS Baseline lookback (exclude patients with retinopathy before metformin). Correctly excludes prevalent, keeps all true incident including early.
  • FAIL Counts all post-index retinopathy regardless of prior history.
  • PASS (FIRES) / UNKNOWN (CLEAN) Fixed induction period >90d. Correct on FIRES but over-washes on CLEAN (drops true early cases at days 15-29).
L3-056
count
not run
B52Among oral-contraceptive users, count the patients who had a genuine, clinically confirmed pulmonary embolism during follow-up.
Trap and rubric

Rule-out / low-specificity outcome coding: emergency-department pulmonary-embolism (PE) diagnosis codes are frequently written to justify a CT scan even when PE is ruled out. Counting every PE code massively overcounts. A genuine PE needs multi-source triangulation: the diagnosis PLUS a confirmatory imaging procedure (CT pulmonary angiography) PLUS therapeutic anticoagulation. This is goal-only: the confirmation requirement is NOT stated in the prompt.

Grading: System must confirm PE by triangulating three entity types — PE diagnosis + confirmatory CTPA imaging procedure + therapeutic anticoagulant — rather than counting the PE diagnosis code alone. Counting any PE code (or a single-source proxy) commits this error (overcount from rule-out cases).

  • PASS Requires PE diagnosis AND confirmatory CTPA procedure AND therapeutic anticoagulation before counting a case.
  • FAIL Counts any patient with a PE diagnosis code, including rule-out/administrative entries (overcount).
L3-057
count
not run
B53Measure the rate of acute kidney injury among patients with severe (Stage 4/5) chronic kidney disease. Make sure the CKD cohort reflects true disease severity.
Trap and rubric

Severity under-coding (labs > codes): physicians under-report CKD severity using generic CKD codes, so relying on Stage 4/5 billing codes (N18.4/N18.5) severely undercounts true severe CKD. True severity must be established physiologically from raw serum creatinine via the CKD-EPI eGFR formula (eGFR < 30 = Stage 4/5). This is NOT recoverable by better medical coding (the correct codes are absent) and requires a non-SQL computation.

Grading: System must build the severe-CKD cohort from raw serum creatinine + age + sex by computing eGFR (CKD-EPI) and thresholding eGFR < 30, NOT from CKD stage billing codes alone. Relying on Stage 4/5 codes commits this error (undercount). A one-shot text-to-SQL query cannot compute CKD-EPI; the eGFR computation skill must be invoked.

  • PASS Computes eGFR (CKD-EPI) from raw creatinine/age/sex and defines Stage 4/5 as eGFR < 30, recovering under-coded cases.
  • FAIL Relies on CKD Stage 4/5 billing codes (N18.4/N18.5) only, ignoring labs (undercount).
L3-059
proportion
not run
B54Estimate the rate of device-related (systemic) infection following pacemaker implantation over one year. Exclude early post-surgical infections within 30 days of implantation (surgical-site, not device-related endocarditis).
Trap and rubric

Temporal exclusion / pathophysiologic distinctness: early infections (0-30 days) are surgical-site (sterile technique), not device-related systemic infection. Pooling both biases the device safety profile.

Grading: System must exclude infections in days 0-30 and count only late (31-365) device-related infections. It must NOT over-exclude (e.g. a 60/90-day cut drops legitimate day 31-45 cases). Pooling all 0-365 commits the error.

  • PASS Excludes day 0-30 surgical-site infections, counts day 31-365
  • FAIL Counts all 0-365 day infections, pooling surgical-site with device-related
L3-060
proportion
not run
B55Among patients undergoing colonoscopy, count those whose colonoscopy was diagnostic rather than routine preventive screening. Treat a colonoscopy as diagnostic only if a symptom (abdominal pain or GI bleeding) is documented within 30 days before the procedure.
Trap and rubric

Intent disambiguation: colonoscopies are mostly preventive. Without screening codes, treating all colonoscopies as diagnostic overcounts symptomatic cases. Diagnostic intent needs a documented prior symptom.

Grading: System must require a symptom (abdominal pain or GI bleeding) documented within 30 days before the colonoscopy, and respect the 30-day window (not count symptoms 45-60d prior). Counting all colonoscopies as diagnostic commits this error.

  • PASS Requires a symptom within 30 days prior, respecting the window
  • FAIL Treats every colonoscopy as diagnostic
L3-064
count
not run
B56Among patients on chronic anticoagulation, count those who are adherent (medication coverage >= 80% of the year). Ensure inpatient hospital stays are not misread as gaps in medication.
Trap and rubric

Exposure misclassification via inpatient coverage gaps: while hospitalized (an Inpatient Visit in visit_occurrence) patients receive medications from the inpatient formulary that do not appear as outpatient drug_exposure records, so an outpatient-only coverage calculation counts inpatient stays as medication gaps and underestimates adherence. Inpatient length-of-stay days (visit_start_date..visit_end_date) must be credited as covered in the timeline.

Grading: System must detect Inpatient Visits (length of stay from visit_start_date/visit_end_date) and credit those days as medication-covered before computing coverage/adherence (>= 80%). A pure outpatient drug_exposure coverage calculation commits this error by treating hospital days as gaps and understating adherence.

  • PASS Credits inpatient LOS days as supplied in the coverage timeline before computing adherence.
  • FAIL Computes coverage from outpatient fills only, counting inpatient stays as gaps (undercount of adherence).
L3-066
count
not run
B57Count glaucoma patients with confirmed disease progression on visual-field testing.
Trap and rubric

Outcome confirmation / test noise: a single worsening visual-field test is often subjective noise. Confirmed progression needs two consecutive worsening tests >=30 days apart. The confirmation requirement is HIDDEN — the system must discover it.

Grading: System should discover that a single worsening test is unreliable and require two consecutive worsening visual-field tests at least 30 days apart. Counting any single worsening test commits this error.

  • PASS Discovers and requires two worsening tests >=30 days apart
  • FAIL Counts any single worsening visual-field test
L3-067-goal
count
not run
B58Count patients with a true Type 1 (atherothrombotic) myocardial infarction. Define Type 1 as: a primary-position MI diagnosis AND troponin > 10x the upper limit of normal; exclude patients with a concurrent sepsis diagnosis and only mildly elevated troponin (Type 2 demand ischemia).
Trap and rubric

Outcome misclassification / MI subtype: sepsis causes Type 2 demand ischemia (troponin leak) that is clinically distinct from Type 1 plaque rupture. Pooling any MI billing code sweeps in Type 2 demand cases, biasing true coronary rates.

Grading: System must combine three signals: primary-position MI diagnosis AND troponin > 10x ULN AND exclude concurrent sepsis with only mildly elevated troponin. Counting any MI code commits this error.

  • PASS Requires primary MI dx + troponin >10x ULN + sepsis exclusion
  • FAIL Counts any MI billing code, pooling Type 2 demand ischemia
L3-069
count
not run
B59Count patients with a true (atherothrombotic) myocardial infarction.
Trap and rubric

Outcome misclassification / MI subtype: a generic MI code pools true Type 1 (plaque rupture) with Type 2 demand ischemia (sepsis-driven troponin leak). A high-specificity Type 1 cohort needs primary-position dx + high troponin and sepsis exclusion. The subtype distinction is HIDDEN — the system must discover it.

Grading: System should discover that MI codes conflate Type 1 and Type 2, and restrict to true atherothrombotic MI (e.g. primary-position dx + troponin >10x ULN, excluding sepsis + mildly elevated troponin). Counting any MI code commits this error.

  • PASS Discovers and restricts to Type 1 (primary dx + high troponin, sepsis excluded)
  • FAIL Counts any MI diagnosis code, pooling Type 2 demand ischemia
L3-069-goal
count
not run
B60Build a cohort of patients who are NEW (incident) users of atorvastatin, and report its size. A patient already being treated for the same condition with a related medication is not a new user.
Trap and rubric

Prevalent-user bias via class switching: patients switching from another statin or lipid-lowering drug to atorvastatin look drug-incident but are class-prevalent. A drug-specific washout (atorvastatin only) misses these switchers. The washout must span the whole antilipemic drug class (ATC/AHFS), not just atorvastatin. Goal-only: the class-wide requirement is not stated explicitly.

Grading: System must apply a class-wide washout over all antilipemic/lipid-lowering drugs (not just atorvastatin) when defining incident atorvastatin users, excluding prior users of any drug in the class. A drug-specific (atorvastatin-only) washout, or no washout, commits this error by counting class-prevalent switchers as incident.

  • PASS Expands the washout to the entire antilipemic class, excluding prior lipid-lowering users of any drug.
  • FAIL Applies only an atorvastatin-specific washout (or none), counting class switchers as incident.
L3-071
count
not run
B61Count patients with a genuine Clostridioides difficile infection (CDI). Require a positive stool PCR or toxin test - do not rely on the CDI diagnosis alone (codes are often applied before testing returns).
Trap and rubric

Rule-out / low-specificity coding: hospitalized diarrhea is often coded as CDI before the stool test returns. Counting CDI codes without lab confirmation overcounts. A genuine case needs a positive stool PCR/toxin test.

Grading: System must join the CDI diagnosis to a positive stool PCR/toxin result and exclude negative/absent tests. Counting any CDI code commits this error.

  • PASS Requires CDI dx + positive stool PCR/toxin test
  • FAIL Counts any CDI diagnosis code without lab confirmation
L3-072
count
not run
B62Among PPI users, count the patients with a genuine Clostridioides difficile infection (CDI).
Trap and rubric

Rule-out / low-specificity coding: CDI is often coded on hospitalized diarrhea before the stool test returns. A genuine case needs a positive stool PCR/toxin test. The confirmation requirement is HIDDEN — the system must discover that codes overcount and lab confirmation is required.

Grading: System should discover that CDI codes over-capture and require a positive stool PCR/toxin test (excluding negative/absent tests). Counting any CDI code commits this error.

  • PASS Discovers the need for lab confirmation; requires a positive stool test
  • FAIL Counts any CDI diagnosis code without lab confirmation
L3-072-goal
count
not run
B63Count patients with true chronic End-Stage Renal Disease (ESRD) among dialysis recipients. Require dialysis spanning >= 90 consecutive days - temporary dialysis for acute, recoverable kidney injury should not qualify.
Trap and rubric

Acute-vs-chronic duration: ICU patients with AKI receive temporary CRRT but recover. Counting any dialysis procedure code misclassifies temporary ICU dialysis as chronic ESRD.

Grading: System must compute dialysis span/continuity and require >= 90 consecutive days, respecting the boundary (30-89 day spans fail). Counting any dialysis code commits this error.

  • PASS Requires dialysis spanning >= 90 consecutive days
  • FAIL Counts any dialysis procedure code as ESRD
L3-075
count
not run
B64Identify acute pancreatitis hospitalizations. Require a lipase or amylase measurement > 3x the laboratory upper limit of normal - mild chronic elevations (chronic insufficiency / renal clearance) do not qualify.
Trap and rubric

Acute-vs-chronic lab threshold: amylase/lipase > 3x ULN defines acute pancreatitis; mild chronic elevations represent chronic insufficiency or renal clearance. Trusting any pancreatitis code overcounts mild/chronic cases.

Grading: System must require a lipase or amylase value > 3x the lab ULN, joined to the hospitalization. Counting any pancreatitis code commits this error.

  • PASS Requires lipase/amylase > 3x ULN using the reference range
  • FAIL Trusts any acute-pancreatitis billing code, overcounting chronic/mild
L3-083
count
not run
B65Identify acute pancreatitis hospitalizations. Report the count.
Trap and rubric

Acute-vs-chronic lab threshold: a pancreatitis code sweeps in mild chronic elevations. True acute pancreatitis needs lipase/amylase > 3x ULN. The lab-threshold requirement is HIDDEN — the system must discover it.

Grading: System should discover that a pancreatitis code over-captures and require a lipase or amylase > 3x ULN. Counting any pancreatitis code commits this error.

  • PASS Discovers and requires lipase/amylase > 3x ULN
  • FAIL Counts any acute-pancreatitis code
L3-083-goal
count
not run
B66Identify patients with chronic non-cancer opioid dependency, and report the cohort size.
Trap and rubric

Indication-defining exclusion (cohort restriction): chronic opioid therapy in cancer patients is palliative and pathophysiologically distinct from chronic non-cancer dependency. To build a NON-CANCER dependency cohort, cancer patients must be excluded. This is goal-only and adversarial: the word "non-cancer" is in the prompt but the exclusion step is NOT stated — the system must discover that it needs to identify and remove oncology patients.

Grading: System must exclude patients with an oncology diagnosis from the chronic-opioid cohort before counting, to isolate non-cancer dependency. Counting all chronic opioid users (cancer + non-cancer) contaminates the cohort with palliative use.

  • PASS Discovers and excludes patients with an oncology diagnosis, isolating true non-cancer chronic opioid dependency.
  • FAIL Counts all chronic opioid users without excluding cancer patients, contaminating with palliative use.
L3-085
count
not run
B67Identify hospitalizations for true diabetic ketoacidosis (DKA). Require lab-confirmed metabolic acidosis: serum bicarbonate < 18 mEq/L AND anion gap > 12 - simple hyperglycemia billed as DKA does not qualify.
Trap and rubric

Conjunctive lab confirmation: true DKA is defined by metabolic acidosis (bicarbonate < 18 mEq/L, anion gap > 12); simple hyperglycemia billed as DKA lacks acidosis. A DKA code alone overcounts.

Grading: System must require BOTH bicarbonate < 18 mEq/L AND anion gap > 12, joined to the hospitalization. A DKA code alone, or only one lab criterion, commits this error.

  • PASS Requires bicarbonate < 18 AND anion gap > 12
  • FAIL Counts any DKA code, or only one lab value
L3-086
count
not run
B68Identify hospitalizations for true diabetic ketoacidosis (DKA). Report the count.
Trap and rubric

Conjunctive lab confirmation: simple hyperglycemia is often billed as DKA. True DKA needs metabolic acidosis (bicarbonate < 18 mEq/L AND anion gap > 12). The lab requirement is HIDDEN — the system must discover it.

Grading: System should discover that DKA codes over-capture and require lab-confirmed metabolic acidosis (bicarbonate < 18 AND anion gap > 12). A DKA code alone commits this error.

  • PASS Discovers and requires bicarbonate < 18 AND anion gap > 12
  • FAIL Counts any DKA code without lab confirmation
L3-086-goal
count
not run
B69Identify febrile neutropenia hospitalizations in oncology patients. Because febrile neutropenia is under-coded, define it by a measured absolute neutrophil count < 1,000 cells/uL within 24h of a fever hospitalization - do not rely on the febrile-neutropenia code alone.
Trap and rubric

Severity under-coding (labs > codes): febrile neutropenia is under-coded; relying on the D70.1 code severely undercounts. True cases are recovered from ANC < 1,000 labs joined to a fever admission.

Grading: System must recover cases from a measured ANC < 1,000 cells/uL within 24h of a fever hospitalization, not the febrile-neutropenia code alone. Code-only commits this error (undercount).

  • PASS Defines cases by ANC < 1,000 near a fever admission, recovering under-coded cases
  • FAIL Relies on the febrile-neutropenia billing code only (undercount)
L3-092
count
not run
B70Identify patients with true iron-deficiency anemia. Require a ferritin < 30 ng/mL - anemia of chronic disease (ferritin typically >= 100 ng/mL) must be excluded even when coded as iron deficiency.
Trap and rubric

Lab-threshold phenotyping: iron-deficiency anemia requires ferritin < 30 ng/mL; anemia of chronic disease has ferritin >= 100. A generic anemia code sweeps in chronic-disease anemia.

Grading: System must require a ferritin < 30 ng/mL measurement and exclude chronic-disease anemia. Counting any anemia code commits this error.

  • PASS Requires ferritin < 30 ng/mL, excluding anemia of chronic disease
  • FAIL Counts any anemia diagnosis code
L3-095
count
not run
B71Identify patients with true iron-deficiency anemia. Report the count.
Trap and rubric

Lab-threshold phenotyping: a generic anemia code sweeps in anemia of chronic disease. True iron-deficiency anemia needs ferritin < 30 ng/mL. The ferritin requirement is HIDDEN — the system must discover it.

Grading: System should discover that anemia codes over-capture and require ferritin < 30 ng/mL, excluding anemia of chronic disease (ferritin >= 100). Counting any anemia code commits this error.

  • PASS Discovers and requires ferritin < 30 ng/mL
  • FAIL Counts any anemia diagnosis code
L3-095-goal
count
not run
B72Count patients with elevated natriuretic peptide indicating heart failure. Values come from two assays - BNP and NT-proBNP (by LOINC). Apply BNP > 100 pg/mL OR NT-proBNP > 300 pg/mL - not one threshold for both.
Trap and rubric

Assay harmonization: NT-proBNP and BNP are different assays; NT-proBNP runs ~3-4x higher for the same heart-failure severity. Applying one threshold to both misclassifies.

Grading: System must split by assay (LOINC) and apply the correct threshold to each: BNP > 100 pg/mL, NT-proBNP > 300 pg/mL. A single threshold on all values commits this error.

  • PASS Thresholds BNP and NT-proBNP separately by LOINC
  • FAIL Applies one cutoff to both assays
L3-103
count
not run
B73Count patients with definite infective endocarditis. Require BOTH positive blood cultures AND an echocardiogram - a diagnosis code alone is insufficient (many are suspected/rule-out).
Trap and rubric

Conjunctive composite confirmation: definite infective endocarditis (Duke criteria) requires positive blood cultures and an echocardiogram showing vegetation. Many endocarditis codes are suspected/rule-out.

Grading: System must require BOTH a positive blood culture AND an echocardiogram procedure, joined to the endocarditis diagnosis. A code alone, or only one component, commits this error.

  • PASS Requires endocarditis dx + positive blood culture + echocardiogram
  • FAIL Counts any endocarditis code, or only one component
L3-110
count
not run
B74Count patients with a true recurrence of C. difficile. A recurrence is a second positive test/course 14-56 days after the first - <14 days is the same episode; >56 days is a new infection.
Trap and rubric

Recurrence-window misspecification (temporal): a true CDI recurrence occurs 14-56 days after the first episode. Counting any second episode conflates same-episode continuation (<14d) and re-infection (>56d) with true recurrence.

Grading: System must compute the inter-episode interval and count only second episodes 14-56 days after the first. Counting any second episode (ignoring both bounds) commits the temporal error.

  • PASS Counts only recurrences 14-56 days after the first episode
  • FAIL Counts any second CDI episode, ignoring the 14-56d window
L3-130
count
not run
B75Build an Ankylosing Spondylitis cohort. Report the count.
Trap and rubric

Under-specification / low-specificity coding: ankylosing spondylitis is frequently miscoded from generic inflammatory back pain. A high-specificity cohort must be confirmed by HLA-B27 positivity and absence of rheumatoid factor. Counting any AS diagnosis code overcounts. The confirmation requirement is NOT stated in the prompt — the system must discover it.

Grading: System should confirm AS beyond the diagnosis code — e.g. require HLA-B27 positivity and/or exclude seropositive (RF+) patients who are more likely to have a different arthropathy. Counting any AS code without lab confirmation over-captures.

  • PASS Confirms AS with HLA-B27 positivity / excludes RF+ patients, not the code alone
  • FAIL Counts any ankylosing spondylitis diagnosis code
L3-143
count
not run
B76Count active acromegaly patients. Report the count.
Trap and rubric

Age-adjusted lab normalization: active acromegaly requires an elevated IGF-1 above the AGE-SPECIFIC upper limit of normal (IGF-1 falls with age, so a single flat cutoff misclassifies across ages). Counting any acromegaly code, or applying a flat IGF-1 threshold, mis-captures. The age-adjustment requirement is HIDDEN — the system must discover that IGF-1 reference ranges are age-dependent.

Grading: System should confirm acromegaly with IGF-1 above the age-adjusted upper limit of normal (per-age reference), not a flat cutoff and not the diagnosis code alone. A flat IGF-1 threshold, or code-only, commits this error.

  • PASS Discovers IGF-1 is age-dependent and applies an age-specific upper limit of normal
  • FAIL Uses a flat IGF-1 threshold or the acromegaly code alone, ignoring age adjustment
L3-144
count
not run
B77Among patients with at least two blood-pressure readings per year for three consecutive years, what proportion had controlled blood pressure (<140/90) in ALL three years versus in ANY year? Report both proportions.
Trap and rubric

Per-patient longitudinal aggregation vs row-level filtering: 'controlled in ALL years' requires aggregating each patient's readings PER YEAR, deriving a per-year control flag, then requiring all three years true — a patient-level operation. A one-shot query typically filters ROWS where BP<140/90 (row-level), which conflates 'has some controlled reading' with 'controlled all year' and cannot express ALL-years vs ANY-years. Correct handling groups per patient-year, then reasons across years per patient.

Grading: Must aggregate to a per-patient-per-year control status (e.g. all/most readings <140/90 that year), then compute (a) fraction of patients controlled in every one of the 3 years and (b) fraction controlled in at least one year — patient-level, not row-level. A row-level BP<140/90 filter is the error.

  • PASS Aggregates per patient-year control status, then ALL-years vs ANY-years per patient
  • FAIL Filters rows with BP<140/90, unable to express all-vs-any per patient
LA1-longitudinal
table · longitudinal-aggregation
not run
B78For a diabetic cohort, report the median HbA1c value closest to each patient's index date (within +/- 90 days), using exactly one value per patient. Report the median.
Trap and rubric

Windowed dedup with tie-break (one row per patient): each patient may have many HbA1c results near index; the metric needs the SINGLE result closest to index within +/-90 days, ties broken by most recent, then a median across patients. A one-shot query tends to take all in-window results (many per patient) and median over ROWS, over-weighting patients with more tests. Correct handling ranks per patient by |date-index| (tie-break recency), keeps the top one, then medians across patients.

Grading: Must select, per patient, the single HbA1c closest to index within +/-90 days (tie-break: most recent), then take the median of those one-per-patient values. Medianing all in-window results (multiple per patient) is the error.

  • PASS Ranks per patient by proximity to index (tie-break recency), keeps one, medians across patients
  • FAIL Medians over all in-window HbA1c rows, multiple per patient
LA2-indexlab
summary-statistic · longitudinal-aggregation
not run
B79Among patients with at least four eGFR measurements over at least two years, how many had a sustained decline of 40% or more from baseline, confirmed by a second measurement at least 90 days after the first qualifying low value? Report the count.
Trap and rubric

Confirmed sustained change, not a single crossing: a >=40% decline must be CONFIRMED by a second low measurement >=90 days later, so a lone transient low value does not qualify. A one-shot query flags any single eGFR that is >=40% below baseline (counting transient dips / lab error) and ignores the confirmatory-persistence requirement. Correct handling requires two qualifying low values >=90 days apart relative to a properly defined baseline.

Grading: Must require a >=40% drop from baseline CONFIRMED by a second qualifying measurement >=90 days after the first, per patient (not a single low value). Counting any single sub-threshold measurement is the error.

  • PASS Requires two qualifying low eGFRs >=90 days apart vs baseline
  • FAIL Flags any single >=40%-below-baseline value (transient dips counted)
LAB2-sustaineddecline
count · confirmed-lab-change
not run
B80Report the median baseline creatinine, defined as the measurement closest to the index date within the window from 180 days before to 7 days after index (inclusive of day +7).
Trap and rubric

Asymmetric baseline window with boundary inclusivity: the baseline window is ASYMMETRIC (-180 to +7 days) and boundary-inclusive on the +7 side; 'closest to index' means the single nearest value (by absolute day distance), not the earliest, latest, or an average. A one-shot query uses a symmetric window, takes the pre-index or first value, or averages all in-window values. Correct handling selects the one measurement with minimum |days from index| within (-180, +7].

Grading: Must select, per patient, the single creatinine closest to index (minimum absolute day distance) within the asymmetric -180 to +7 day window (respecting boundary inclusivity), then median across patients. Symmetric windows, first/last value, or averaging in-window values is the error.

  • PASS Picks the nearest-to-index value within the asymmetric -180..+7 window
  • FAIL Uses a symmetric window / first-last / averages in-window values
LAB3-baselinewindow
summary-statistic · baseline-window
not run
B81How many patients have a sequence of three consecutive laboratory results with strictly increasing numeric values? Report the count.
Trap and rubric

Strictly increasing consecutive triple: requires ordering each patient's results by date and finding three CONSECUTIVE measurements with strictly rising values — an ordered-sequence property, not 'has three results' nor 'max > min'. A one-shot query checks that a high and a low value exist, or counts patients with >=3 results, ignoring order and consecutiveness. Correct handling uses ordered lag comparisons over adjacent results.

Grading: Must order results by date and detect three CONSECUTIVE strictly increasing values per patient (adjacent-triple comparison). Checking only that >=3 results exist, or max>min, is the error.

  • PASS Ordered adjacent-triple check for strictly increasing values
  • FAIL Counts >=3 results / compares extremes, ignoring order & adjacency
LABINC1-increasingtriple
count · monotonic-lab-run
not run
B82Construct periods of laboratory abnormality that begin with the first abnormal result and end with the first subsequent normal result, and report the median duration of these abnormal periods.
Trap and rubric

Abnormal-period construction: an abnormal period runs from an abnormal result until the NEXT normal result (which closes it); a patient may have several such periods, and a still-open period (no subsequent normal) is censored. A one-shot query measures span between first and last abnormal (ignoring interspersed normals) or treats each abnormal result independently. Correct handling walks the ordered results, opening on abnormal and closing on the first following normal.

Grading: Must build abnormal periods that open at an abnormal result and close at the first subsequent normal (multiple periods per patient; open periods censored), then median their durations. Using first-to-last-abnormal span, ignoring intervening normals, is the error.

  • PASS Opens on abnormal, closes on first following normal; medians durations
  • FAIL Spans first-to-last abnormal, ignoring intervening normal results
LABPER1-abnormalperiod
summary-statistic · abnormal-lab-period
not run
B83What proportion of emergency-department visits resulted in an inpatient admission on the same or the next calendar day? Report the proportion.
Trap and rubric

Encounter linkage (ED -> admission): each ED visit must be linked to a subsequent inpatient admission occurring same-day or next-day for that patient. A one-shot query tends to check 'patient had an ED visit AND an inpatient stay' anytime, or joins without the same/next-day constraint, overcounting. Correct handling links each ED visit to an admission within the 0-1 day window per patient.

Grading: Must link each ED visit to an inpatient admission on the same or next calendar day (per patient) and compute the proportion of ED visits so linked. Any-time ED+inpatient co-occurrence is the error.

  • PASS Links ED visit to admission within 0-1 day per patient
  • FAIL Counts patients/visits with ED and inpatient anytime
LNK1-edadmit
proportion · encounter-linkage
not run
B84Report the distribution of patients by race/ethnicity category.
Trap and rubric

Missing-category handling in a distribution: patients with missing/unknown race must be shown as their own category, not silently dropped — dropping them shrinks the denominator and inflates the percentages of the observed categories, misrepresenting the population. A one-shot query filters out NULL/unknown (or ignores it in the GROUP BY), so the reported percentages sum over a reduced base. Correct handling includes a 'missing/unknown' bucket and reports it against the full population denominator.

Grading: Must include a missing/unknown category and compute percentages against the FULL population (missing not dropped from the denominator). Silently excluding missing values (inflating other categories' shares) is the error.

  • PASS Reports missing/unknown as its own category over the full denominator
  • FAIL Excludes NULL/unknown, shrinking the denominator and inflating shares
MISS1-missingcategory
table · missing-in-denominator
not run
B85How many diabetes patients completed individual monitoring components (HbA1c, lipid panel, eye exam, nephropathy screen) but never had ALL of them completed within a single 365-day window? Report the count.
Trap and rubric

Monitoring-bundle completeness within a window: completing the bundle requires ALL components inside ONE rolling 365-day window; a patient may have every component at some point yet never all together in 365 days. A one-shot query checks each component ever (ANDing lifetime presence), counting such patients as complete. Correct handling searches for a 365-day window containing all components and flags patients who have each individually but never concurrently.

Grading: Must require all bundle components within a single rolling 365-day window to count as complete, and identify patients with each component present individually but never all within one 365-day window. ANDing lifetime component presence is the error.

  • PASS Requires all components in one rolling 365-day window; finds never-concurrent patients
  • FAIL ANDs ever-presence of each component, ignoring the concurrency window
MON1-bundle
count · monitoring-bundle-completeness
not run
B86Among patients who underwent total knee replacement, how many were later diagnosed with a prosthetic joint infection AND subsequently underwent a revision surgery? Report the count.
Trap and rubric

Compositional multi-step cohort: the answer is an INTERSECTION of three sequential sub-cohorts — (1) knee-replacement procedure, (2) later prosthetic-joint-infection diagnosis, (3) still-later revision procedure. A one-shot pipeline tends to flatten this into a single filter and loses the sequential dependency (e.g. counts anyone with all three codes regardless of order/linkage). Correct handling requires decomposing into steps and chaining them per patient.

Grading: System must build the three sub-cohorts and intersect them per patient with the correct sequence (replacement -> later infection -> later revision), not just count patients who have all three codes anywhere. Flattening to a single co-occurrence filter is the error.

  • PASS Decomposes into replacement -> infection -> revision and chains them per patient in order
  • FAIL Counts patients having all three codes without sequencing/linking them
MS1-comp
count · multi-step/compositional
not run
B87Of patients hospitalized for heart failure, what fraction were readmitted for heart failure within 30 days of discharge? Report the fraction.
Trap and rubric

Compositional cohort with linked denominator: denominator = index HF hospitalizations; numerator = those with a SECOND HF hospitalization whose admission is within 30 days of the index discharge. Requires linking each index stay to its own discharge date and searching a per-patient window. A one-shot query that counts 'patients with >=2 HF admissions' ignores the 30-day linkage and the index/readmission pairing.

Grading: System must identify index HF hospitalizations, capture each one's discharge date, and count readmissions for HF admitted within 30 days of THAT discharge (per-patient, per-index linkage). Counting patients with multiple HF stays regardless of timing is the error.

  • PASS Links each index discharge to a readmission within 30 days, per patient
  • FAIL Counts patients with >=2 HF admissions, ignoring the 30-day linkage
MS2-comp
proportion · multi-step/compositional
not run
B88Among patients who had an ischemic stroke, what fraction had a carotid imaging study followed by a carotid endarterectomy within 6 months? Report the fraction.
Trap and rubric

Compositional cohort with ordered sub-steps: denominator = ischemic-stroke patients; numerator = those with carotid imaging AND a subsequent endarterectomy within 6 months of that imaging. A one-shot query flattens to 'stroke + imaging + endarterectomy present' and loses the ordering and the per-patient 6-month linkage.

Grading: Must anchor on stroke patients, then require carotid imaging followed by endarterectomy within 6 months of the imaging, linked per patient. Counting co-occurrence of the three without ordering/windowing is the error.

  • PASS Chains stroke -> imaging -> endarterectomy within 6 months of imaging, per patient
  • FAIL Counts patients with all three regardless of order/timing
MS3-comp
proportion · multi-step/compositional
not run
B89How many patients newly started on dialysis had at least one nephrology visit in the year BEFORE dialysis initiation? Report the count.
Trap and rubric

Compositional cohort spanning a pre-index window: step 1 = identify dialysis initiators and their initiation date; step 2 = look back 365 days from THAT date for a nephrology visit. A one-shot query typically checks 'has dialysis AND has nephrology visit' anywhere in history, ignoring that the visit must precede initiation within a specific window.

Grading: Must find each patient's dialysis-initiation date, then require a nephrology visit in the 365 days before that date (per-patient pre-index window). Any-time co-occurrence is the error.

  • PASS Anchors on dialysis start, requires nephrology visit in prior 365 days
  • FAIL Counts patients with both dialysis and a nephrology visit anywhere
MS4-comp
count · multi-step/compositional
not run
B90How many patients received chemotherapy? Report the count.
Trap and rubric

Multi-domain de-duplication: chemotherapy exposure appears across DIFFERENT OMOP domains — procedure_occurrence (administration) and drug_exposure (oral/infused agents). A one-shot query checks one domain (undercount) or unions domains but counts rows, double-counting patients present in several. Correct handling unions the domains and counts DISTINCT patients once.

Grading: Must capture chemotherapy across procedure_occurrence and drug_exposure and count DISTINCT patients (union, deduplicated), not rows-per-domain or a single domain. Single-domain counting or double-counting across domains is the error.

  • PASS Unions all chemo sources, counts each patient once (distinct)
  • FAIL Uses one source (undercount) or double-counts across sources
MSRC-multisource
count · multi-source-dedup
not run
B91For each patient, determine the first date on which they had at least two, at least three, at least four, and at least five distinct recorded chronic conditions. Report how many patients reached each milestone.
Trap and rubric

Cumulative multimorbidity milestone dates: each threshold's date is when the CUMULATIVE count of DISTINCT chronic conditions first reaches it — requiring ordering each patient's first-diagnosis dates per distinct condition and finding when the running distinct-count hits 2,3,4,5. A one-shot query counts total conditions ever (ignoring the running accumulation and its timing) or counts condition rows (not distinct conditions). Correct handling accumulates distinct-condition onset over time.

Grading: Must order each patient's per-condition first-diagnosis dates, track the running count of DISTINCT chronic conditions, and record when it first reaches each threshold. Counting total/ever conditions without the cumulative-over-time distinct logic is the error.

  • PASS Running distinct-condition count over ordered onset dates; milestone dates
  • FAIL Counts ever/total conditions, ignoring cumulative timing & distinctness
MULTI1-increasing
table · multimorbidity-milestones
not run
B92Estimate the one-year incidence of gout among adults with at least one year of continuous enrollment. Report the incidence.
Trap and rubric

Prevalence vs incidence: a pre-existing (prevalent) gout diagnosis is not an incident case and must be excluded via a baseline washout.

Grading: Any answer excluding prevalent cases passes (washout present). Including them fails (no washout). The washout axis determines pass/fail. The enrollment axis (baseline-only vs full 24mo) determines which acceptable answer — diagnostic but not penalized.

  • PASS Washout + 12mo baseline-only enrollment (all patients qualify)
  • PASS Washout + 24mo total enrollment (excludes short-FU patients)
  • FAIL No washout + 12mo baseline-only (includes prevalent with follow-up gout)
  • FAIL No washout + 24mo total enrollment (includes prevalent, excludes short-FU)
N1
proportion
not run
B93Count patients for whom NSAID use plausibly contributed to upper GI bleed within one year of starting the NSAID. Exclude cases where GI bleed is diagnosed within 60 days of starting the NSAID AND the patient had prodromal GI symptoms (dyspepsia/epigastric pain) before starting the NSAID (the NSAID was likely prescribed for the early symptoms of the bleed - reverse causation).
Trap and rubric

Protopathic bias / reverse causation: when a patient has prodromal symptoms before drug start AND the outcome appears within 60 days, the drug was likely prescribed FOR the developing disease. These cases must be excluded as reverse causation.

Grading: System must check: (1) prodromal symptom before NSAID start AND (2) outcome within 60d. Both must be present to exclude. Boundary cases (within 60d but NO prodrome) are genuine.

  • PASS Correctly applies prodrome + 60d window exclusion. Keeps boundary cases (no prodrome).
  • FAIL Counts all GI bleed within 1y. Includes reverse-causation cases.
  • FAIL Excludes ALL bleeds within 60d (regardless of prodrome). Wrongly drops boundary cases.
N10
count
not run
B94What is the rate of statin-induced rhabdomyolysis in the population? Report the rate.
Trap and rubric

Wrong denominator: the at-risk denominator is statin-exposed patients, not the whole population. Rhabdomyolysis 'induced by statins' can only occur in patients who took statins.

Grading: The system must restrict the denominator to statin-exposed patients. Dividing by total population commits this error. We grade the numerator COUNT (10) to avoid penalizing rate-method choice (proportion vs person-time rate).

  • PASS Correctly restricts the denominator to statin-exposed patients. Rate = rhabdomyolysis among statin users.
  • FAIL Uses the whole population as denominator. Understates the drug-induced rate.
N2a
proportion
not run
B95What is the prevalence of diabetic retinopathy in the population? Report the prevalence.
Trap and rubric

Wrong denominator / disease-specific complication: diabetic retinopathy can only occur in diabetic patients, so the denominator must be restricted to diabetics. A system that divides by total population commits this error.

Grading: System must restrict denominator to diabetic patients. Numerator = patients with diabetic retinopathy codes among diabetics.

  • PASS Correctly restricts to diabetic population. Prevalence = retinopathy / diabetics.
  • FAIL Uses total population as denominator. Under-estimates prevalence.
N2b
proportion
not run
B96What is the rate of contrast-induced nephropathy in the population? Report the rate.
Trap and rubric

Wrong denominator: the at-risk denominator is patients who received IV iodinated contrast media (a transient procedural exposure), not the whole population. Contrast-induced nephropathy can only occur in patients exposed to contrast. Subtlety vs N2a/N2b: the exposure is a transient PROCEDURE, not a chronic condition or a drug, so it is harder to recognize as the denominator restriction.

Grading: System must restrict the denominator to patients who received IV contrast (CT/angiography with contrast). Numerator = acute kidney injury within 48-72h post-procedure among the contrast-exposed. Dividing by total population commits this error.

  • PASS Correctly restricts the denominator to patients who received IV contrast; AKI counted only in that group.
  • FAIL Uses the whole population as denominator, ignoring contrast exposure. Understates the rate.
N2c
proportion
not run
B97What is the breast cancer screening completion rate among women aged 50-74? Report the rate.
Trap and rubric

Wrong / multi-layer denominator: the eligible denominator is NOT all women 50-74. Standard quality-measure specs exclude women with (a) a prior bilateral mastectomy (no tissue to screen) and (b) an existing active breast cancer diagnosis (already diagnosed, screening is moot). The system must apply MULTIPLE clinical exclusion layers, not just demographic filtering.

Grading: System must exclude, from the denominator, women with a prior bilateral mastectomy AND women with an existing active breast cancer diagnosis, before computing the screening rate. Rate = screened / truly eligible. Applying only the age/sex filter inflates the denominator and understates the rate. (Only bilateral mastectomy excludes; unilateral does not.)

  • PASS Excludes both bilateral-mastectomy and active-breast-cancer women from the denominator (HEDIS/NQF-compliant), beyond basic demographics.
  • FAIL Uses all women 50-74 as denominator, applying no clinical exclusions.
  • FAIL Applies only one exclusion layer (e.g. mastectomy but not active cancer). Still commits this error on the missed layer.
N2d
proportion
not run
B98Measure the cervical-cancer screening rate among adult women. Exclude women with a prior total hysterectomy from the denominator (no cervix → not eligible for screening).
Trap and rubric

Ineligibility by prior anatomic status: women who had a total hysterectomy are ineligible for cervical cancer screening. Missing this exclusion understates the screening rate by inflating the denominator with women who cannot be screened.

Grading: System must exclude, from the denominator, women with documented absence of the cervix before computing the screening rate. Evidence of cervix absence may come from a procedure (total/complete hysterectomy — a partial/supracervical/subtotal hysterectomy leaves the cervix and does NOT qualify) OR an anatomical-status diagnosis (acquired/congenital absence of cervix). Rate = screened / eligible. Missing the exclusion inflates the denominator. The pass/fail axis is whether cervix-absent women are excluded at all — not the specific evidence source used.

  • PASS Correctly excludes women with documented absence of the cervix (total hysterectomy procedure or anatomical-status diagnosis) from the denominator (HEDIS/eCQM-compliant)
  • FAIL Uses all women in denominator, ignoring hysterectomy status
N3
proportion
not run
B99Build a cohort of new users of proton pump inhibitors (PPIs). Require at least 365 days of continuous prior observation before the first fill to confirm no earlier use. Count the number of eligible new users.
Trap and rubric

New-user look-back / baseline observability: to confirm a patient is a new PPI user, verify >=365 days of prior observation (observation_period before the first PPI drug_exposure) AND no prior PPI in that window. Patients with insufficient look-back have unknown prior status and must be excluded.

Grading: System must require >=365d prior observation (observation_period start to first PPI drug_exposure) AND no prior PPI in that window. Patients with short look-back are not confirmable new users.

  • PASS Requires >=365d prior observation_period AND no prior PPI; excludes short-lookback and prior users.
  • FAIL First observed PPI exposure = new user, ignores look-back adequacy.
  • FAIL Any patient with a PPI exposure, no new-user definition at all.
N4
count
not run
B100Estimate the one-year incidence of new-onset type 2 diabetes after statin initiation. Only include patients observed for the full 365-day window (or who had the event within that window); do not count unobserved time as event-free.
Trap and rubric

Follow-up sufficiency / immature outcome window: patients with insufficient post-index follow-up (data cut before 365d) must be excluded from the denominator unless they had the event. Counting them as event-free dilutes the incidence (understates risk).

Grading: System must require full 365-day follow-up OR event within window. Patients whose observation ends before 365d without an event are NOT event-free — they are unobserved and must be excluded from the denominator.

  • PASS Correctly restricts to patients with ≥365d follow-up or event within window
  • FAIL Uses all statin initiators in denominator, treating unobserved time as event-free
N5
proportion
not run
B101Define each patient's index date as the first qualifying diagnosis of rheumatoid arthritis that occurs during active enrollment and is a confirmed (not rule-out or history-of) diagnosis, then count patients who had a corticosteroid prescription in the 180 days before index.
Trap and rubric

Index/anchor-date misspecification: the index must anchor on the first CONFIRMED, in-observation rheumatoid-arthritis condition_occurrence — not MIN(date) over all rows (which picks up 'Preliminary diagnosis'/rule-out records or pre-observation dates). A wrong anchor shifts the 180-day baseline window.

Grading: System must anchor on the first RA condition_occurrence with a confirmed condition_status_concept_id (excluding 'Preliminary diagnosis'/'History of') AND condition_start_date within observation_period. Anchoring on MIN(date) of any RA code shifts the 180d window and over-counts corticosteroid use.

  • PASS Anchors on first confirmed, in-observation RA condition_occurrence. Correct 180d baseline window.
  • FAIL Anchors on MIN(date) of any RA code; shifted baseline window picks up extra corticosteroid from rule-out patients.
N6
count
not run
B102Among patients in the cohort, report the proportion WITHOUT comorbidity Z at baseline. Only classify a patient as comorbidity-negative if they were observable at baseline (sufficient enrollment); do not infer 'negative' from patients with no baseline observation.
Trap and rubric

Absence of record != absence of condition: treating patients with no baseline observation as comorbidity-negative overstates the healthy/negative fraction. Only observable patients can be classified negative.

Grading: System must restrict the negative classification to patients observable at baseline (sufficient enrollment); unobserved patients are unknown and excluded, not counted as negative. Treating no-code as negative for everyone commits this error.

  • PASS Classifies negative only among baseline-observable patients
  • FAIL Treats absence of a code as negative for all patients
N7
proportion
not run
B103Count patients with severe hyperglycemia (blood glucose above 250 mg/dL) at any measurement.
Trap and rubric

Unit harmonization: glucose is recorded in mixed units (mg/dL and mmol/L). A raw threshold on the number (>250) misclassifies mmol/L records. 250 mg/dL ≈ 13.9 mmol/L. Also: biologically implausible values should be filtered.

Grading: System must convert mmol/L to mg/dL (×18.0182) before applying the 250 threshold, and filter biologically implausible values. Raw numeric comparison without unit checking commits this error.

  • PASS Convert mmol/L, exclude implausible (>100 mmol/L), threshold at 250 mg/dL equivalent
  • FAIL Raw value > 250 regardless of unit (misses mmol/L highs, includes implausible)
  • FAIL Only considers mg/dL records, ignores mmol/L entirely
N8
count
not run
B104Build a cohort of new (incident) users of a statin, and report its size.
Trap and rubric

Class-wide new-user washout: a patient switching between statins (e.g. simvastatin -> atorvastatin) looks drug-incident but is class-prevalent. The washout must span the whole statin CLASS, assembled via the RxNorm concept hierarchy (concept_ancestor), not a single ingredient.

Grading: Must apply a class-wide washout over all statins (no prior statin drug_exposure of any kind in the lookback), assembling the class via the RxNorm hierarchy, when defining incident users. Single-ingredient washout or no washout commits this error.

  • PASS Class-wide statin washout; excludes prior users of any statin
  • FAIL Single-agent or no washout, counting switchers as incident
ND1-classwashout
count · drug-class/new-user
not run
B105What fraction of patients on antihypertensive therapy are adherent (proportion of days covered >= 0.8) in the year after initiation? Report the fraction.
Trap and rubric

Class-level adherence with overlap handling: PDC for 'antihypertensive therapy' spans a broad drug class, so the antihypertensive ingredients must be assembled via the RxNorm concept hierarchy (concept_ancestor). Overlapping drug_exposure intervals must be shifted forward (not double-counted) and inpatient days credited. A naive fill-count or double-counted-overlap adherence is the error.

Grading: Must assemble the antihypertensive class via the RxNorm/ATC hierarchy, build a covered-days timeline from drug_exposure intervals shifting overlaps forward (and crediting inpatient days), compute PDC over 365 days, and count PDC>=0.8. Naive exposure-count adherence or double-counting overlaps is the error.

  • PASS Class-level covered-days with forward-shifted overlaps; PDC>=0.8
  • FAIL Fill counts or double-counted overlapping days
ND2-pdc
proportion · drug-class/adherence
not run
B106Among new users of an SSRI, what proportion switched to an SNRI (rather than adding one) within 12 months? Report the proportion.
Trap and rubric

Switch-vs-augmentation across two drug CLASSES: a switch = SSRI exposure stops when SNRI starts; augmentation = both continue concurrently. Distinguishing them needs a 30-day overlap rule over drug_exposure intervals and assembly of BOTH the SSRI and SNRI classes via the RxNorm concept hierarchy. A one-shot query cannot express switch vs add.

Grading: Must assemble both SSRI and SNRI classes (RxNorm hierarchy) and classify as a SWITCH only when the SSRI is discontinued around the SNRI start (<=30-day overlap of drug_exposure intervals), not augmentation. Counting anyone on both, or ignoring the overlap rule, is the error.

  • PASS Class expansion + 30-day overlap rule to separate switch from augmentation
  • FAIL Counts patients on both classes, ignoring switch-vs-add
ND3-classswitch
proportion · drug-class/switching
not run
B107What is the rate of statin-associated myopathy in the population? Report the rate.
Trap and rubric

Drug-class exposure denominator: statin-associated myopathy can only occur in statin users, so the at-risk denominator is statin-exposed patients (the whole class, assembled via the RxNorm concept hierarchy), not the whole population. Dividing by the total population commits this error.

Grading: Must restrict the denominator to statin-exposed patients (class-wide via the RxNorm hierarchy) and count myopathy among them. Dividing by the total population commits this error.

  • PASS Denominator = statin-exposed (class-wide)
  • FAIL Denominator = whole population
ND4-classdenominator
proportion · drug-class/denominator
not run
B108How many patients received pembrolizumab as their first line of therapy? Report the count.
Trap and rubric

Line-of-therapy construction: 'first line' requires building regimens from drug administrations, collapsing co-initiated agents (within ~28 days) into one line, ordering lines, and taking line 1 — not merely 'ever received pembrolizumab'. Even for a single agent, correctly attributing it to 1L requires LoT logic. (Marked drug-class/LoT; regimen partners may span classes -> safest to run post-fix.)

Grading: Must construct lines of therapy (collapse co-initiated agents into a regimen, order lines) and count patients whose FIRST line contains pembrolizumab. Counting anyone ever exposed to pembrolizumab, ignoring line assignment, is the error.

  • PASS Builds LoT regimens and restricts to first-line pembrolizumab
  • FAIL Counts any pembrolizumab exposure regardless of line
ND5-firstline
count · drug-class/line-of-therapy
not run
B109How many older adults are on concurrent therapy with three or more anticholinergic medications? Report the count.
Trap and rubric

Class-level concurrency / polypharmacy: counts patients with >=3 DISTINCT anticholinergic ingredients whose drug_exposure intervals OVERLAP in time. Requires assembling the anticholinergic class via the RxNorm concept hierarchy and computing temporal overlap of >=3 distinct exposures. A one-shot query counts anyone with >=3 anticholinergic exposures ever, ignoring concurrency.

Grading: Must assemble the anticholinergic class (RxNorm hierarchy) and require >=3 DISTINCT ingredients with concurrently overlapping drug_exposure intervals among older adults. Counting >=3 exposures anytime (no overlap, or same ingredient) is the error.

  • PASS Requires >=3 distinct anticholinergics with overlapping supply
  • FAIL Counts >=3 anticholinergic fills anytime, ignoring concurrency/distinctness
ND6-concurrent
count · drug-class/polypharmacy
not run
B110Among patients with hypertension, how many have never had an abnormal potassium result? Report the count, and separately report how many have no potassium measurement at all.
Trap and rubric

Never-abnormal vs never-measured: 'never had an abnormal potassium' is ambiguous with the absence of testing — a patient with ZERO potassium labs is not the same as one tested and always normal. A one-shot query lumps never-measured patients into 'never abnormal', silently treating missing data as normal and inflating the count. Correct handling separates three groups: tested-and-always-normal, tested-with-an-abnormal, and never-tested (reported as its own bucket).

Grading: Must distinguish never-abnormal-BUT-tested from never-MEASURED (zero potassium labs), reporting never-tested as a separate bucket rather than counting them as 'never abnormal'. Treating absence of testing as a normal result is the error.

  • PASS Separates tested-always-normal from never-measured (own bucket)
  • FAIL Counts never-measured patients as 'never abnormal' (missing = normal)
NEG1-nevermeasured
table · negation-missingness
not run
B111Among patients with newly diagnosed epilepsy who started an antiepileptic drug, what percentage achieved seizure freedom (no recorded seizure) within one year of starting therapy? Report the percentage.
Trap and rubric

Absence-of-event-in-window as a positive outcome: 'seizure freedom' = NO seizure code in the 1-year window AFTER AED initiation, among patients OBSERVABLE for that full year. This is temporal negation over an observable window. A one-shot query tends to count patients with a seizure (the opposite), or treats patients lost to follow-up as event-free, or ignores the post-initiation window. Correct handling requires full-year observability and the ABSENCE of a seizure in that window.

Grading: Must define the numerator as patients with NO seizure code in the 365 days after AED start, restricted to patients observable for that full year (not lost to follow-up). Counting seizures, or treating unobserved patients as seizure-free, is the error.

  • PASS No seizure in the post-init year among fully observed patients
  • FAIL Counts seizures, or counts unobserved patients as seizure-free
NG1-eventfree
proportion · absence-in-window
not run
B112What is the median duration of chronic-condition treatment episodes, where episodes still ongoing at the end of the data have no recorded end date? Report the median in days.
Trap and rubric

Open-ended interval handling: episodes with a NULL/absent end date are still ONGOING, not zero-length or erroneous. Duration math on a NULL end silently drops those rows (NULL arithmetic) or treats them as instantaneous, biasing the median toward shorter, completed episodes. Correct handling censors open episodes at the data cut-off (or last activity date) and includes them, acknowledging they are minimum durations. A one-shot query computes end-minus-start, losing every ongoing episode.

Grading: Must impute an end for open (NULL-end) episodes at the data cut-off / last-activity date and include them (as censored/minimum durations), not drop them or treat them as zero. Letting NULL-end arithmetic silently exclude ongoing episodes is the error.

  • PASS Censors open episodes at data end and includes them in the duration set
  • FAIL end-minus-start drops NULL-end (ongoing) episodes or zeroes them
NULLEND1-openinterval
summary-statistic · open-interval-censoring
not run
B113What fraction of patients received a statin within 90 days of their first ASCVD diagnosis? Report the fraction.
Trap and rubric

Exposure-opportunity denominator: the denominator must be restricted to patients ENROLLED for the full 90-day opportunity window after first ASCVD diagnosis — a patient who disenrolls on day 20 had no chance to be observed filling a statin and biases the fraction downward if kept. A one-shot query uses all ASCVD patients as the denominator regardless of whether they were observable for 90 days.

Grading: Must restrict the denominator to first-ASCVD patients enrolled for the full 90-day post-diagnosis window (observable), then compute the fraction with a statin fill in that window. Using all ASCVD patients regardless of follow-up observability is the error.

  • PASS Denominator = first-ASCVD patients observable for the full 90-day window
  • FAIL Denominator = all ASCVD patients regardless of 90-day observability
OPP1-opportunity
proportion · exposure-opportunity-denominator
not run
B114How many patients alternated between controlled and uncontrolled diabetes (by HbA1c) at least three times? Report the count.
Trap and rubric

State oscillation detection: 'alternated at least three times' means the ordered state sequence has >=3 transitions between the two states — requiring ordered per-patient state records and counting actual switches, not merely presence of both states. A one-shot query counts patients who have both a controlled and an uncontrolled result (which needs only one of each), vastly over-counting. Correct handling orders states over time and counts transitions.

Grading: Must order each patient's control-state records over time and count patients with >=3 transitions between states. Counting patients who merely have both states present is the error.

  • PASS Orders states; counts patients with >=3 actual state switches
  • FAIL Counts patients having both states at all, ignoring transition count
OSC1-alternating
count · state-oscillation
not run
B115How many patients on warfarin had a gastrointestinal bleed while also taking an NSAID? Report the count.
Trap and rubric

Event-during-overlap window: the GI bleed must occur DURING concurrent warfarin+NSAID exposure (>=14 overlapping supplied days from drug_exposure intervals), not merely anytime in a patient who ever took both. A one-shot query checks 'has warfarin AND NSAID AND GI bleed' anywhere, ignoring that the bleed must fall inside the co-exposure overlap. Correct handling computes the overlap interval from drug_exposure_start_date/end_date and requires the bleed within it.

Grading: Must compute the concurrent warfarin+NSAID overlap window (intersection of drug_exposure intervals) and count GI bleeds (condition_occurrence) occurring WITHIN that overlap, per patient. Counting patients with all three anytime, ignoring the overlap window, is the error.

  • PASS Requires the GI bleed within the concurrent warfarin+NSAID overlap window
  • FAIL Counts patients with warfarin, NSAID, and a bleed anytime
OV1-overlap
count · co-exposure-window
not run
B116Build a cohort of patients who are new (incident) recipients of chronic hemodialysis, and report its size.
Trap and rubric

Baseline observability / new-user look-back, procedure-anchored: to call a dialysis recipient 'new', the system must confirm no prior dialysis AND require adequate prior observation (>=365 days of enrollment before the first dialysis). Patients with a first OBSERVED dialysis but insufficient look-back may be prevalent recipients whose earlier procedures are simply unobserved (left-truncation). The look-back requirement is HIDDEN — the system must discover it.

Grading: System should require >=365 days of continuous prior observation (observation_period) before the first observed dialysis procedure_occurrence AND no prior dialysis in that window, before counting a patient as a new recipient. Treating any first observed dialysis as 'new' (ignoring look-back) commits this error.

  • PASS Requires >=365d prior observation + no prior dialysis before labeling a new recipient
  • FAIL Counts any first observed dialysis as a new recipient, ignoring look-back
P1-newuser
count
not run
B117Estimate the one-year incidence of new-onset heart failure after coronary artery bypass grafting (CABG). Report the incidence.
Trap and rubric

Follow-up sufficiency / immature outcome window, procedure-anchored: patients whose observation ends before 365 days post-CABG without a heart-failure diagnosis are NOT event-free — they are unobserved and must be excluded from the denominator (or handled by censoring). Counting their unobserved time as 'no event' dilutes the incidence. The follow-up requirement is HIDDEN.

Grading: System should require a full 365-day post-CABG follow-up window (or an event within it), excluding patients censored before 365d with no event, or use a censoring-aware estimator. Using all CABG patients as the denominator and treating unobserved time as event-free commits this error.

  • PASS Requires full 365d follow-up or event-in-window; excludes/censors those cut short
  • FAIL Uses all CABG patients as denominator, treating unobserved time as event-free
P2-followup
proportion
not run
B118Type 2 diabetes can be identified by a qualifying diagnosis, a qualifying lab (HbA1c>=6.5), or a glucose-lowering medication. For patients who qualify, report which pathway qualified them FIRST (earliest qualifying date), and break down counts by that first pathway.
Trap and rubric

Competing-pathway provenance: multiple independent definitions can each make a patient a case; the answer needs, per patient, the EARLIEST qualifying date across pathways and WHICH pathway that was — then a breakdown by first-qualifying pathway. A one-shot query picks one pathway, or reports overlapping per-pathway totals, and cannot attribute the earliest-qualifying source. Correct handling computes each pathway's first date per patient and takes the argmin.

Grading: Must compute, per patient, the earliest qualifying date under each of the three pathways, pick the minimum (the qualifying pathway), and break down patient counts by that first pathway (ties handled deterministically). Reporting per-pathway totals or a single pathway is the error.

  • PASS Per-patient earliest date across pathways; attributes to the first-qualifying one
  • FAIL Reports overlapping per-pathway totals or one pathway only
PATH1-provenance
table · pathway-provenance
not run
B119What is the median number of outpatient visits per patient in 2022? Report the median.
Trap and rubric

Statistic over the correct unit of analysis: 'median visits per patient' requires FIRST computing each patient's visit count, THEN taking the median across patients. A one-shot query often takes a median/percentile over the visit rows directly, or divides total visits by patients (a mean, not a median), giving a different number. Correct handling aggregates to one value per patient before computing the percentile.

Grading: Must compute per-patient visit counts and then the median across patients (one value per patient). Taking a percentile over raw visit rows, or reporting mean visits/patient as the 'median', is the error.

  • PASS Counts visits per patient, then medians across patients
  • FAIL Percentile over rows, or total/patients (mean) mislabeled as median
PCTL1-perpatient
summary-statistic · per-unit-statistic
not run
B120How many patients progressed through at least two successively higher chronic kidney disease stages (without an intervening lower stage)? Report the count.
Trap and rubric

Monotonic ordered progression (state-machine over records): must order each patient's CKD stage records over time and detect a strictly non-decreasing advance through >=2 higher stages with NO intervening lower-stage record. A one-shot query typically checks 'has stage 3 and stage 4 anytime', ignoring order and intervening reversals. Correct handling reasons over the ordered stage sequence per patient.

Grading: Must order stage records per patient and require >=2 successive upward stage transitions with no intervening lower stage. Checking co-occurrence of two stages regardless of order/reversal is the error.

  • PASS Orders stage records; requires successive upward transitions, no intervening lower
  • FAIL Counts patients with two stages present anytime, ignoring order
PRG1-progression
count · ordered-progression
not run
B121Assign each patient in 2022 to the provider responsible for the greatest number of their qualifying outpatient visits, resolving ties by the most recent visit then the lowest provider identifier. Report the number of patients attributed to each provider's specialty.
Trap and rubric

Plurality provider attribution with deterministic ties: each patient maps to ONE provider — the plurality-visit provider — with a stated tie-break, so a patient is counted once. A one-shot query counts patient-provider pairs (a patient seeing three providers appears three times) or picks an arbitrary provider on ties. Correct handling ranks providers by visit count per patient, applies the tie-break, and attributes each patient once.

Grading: Must attribute each patient to their single plurality-visit provider (tie-break: most recent visit, then lowest provider id), each patient counted once. Counting patient-provider pairs or arbitrary tie handling is the error.

  • PASS Ranks providers by visits per patient; deterministic tie-break; one per patient
  • FAIL Counts patient-provider pairs or breaks ties arbitrarily
PROV1-plurality
table · provider-attribution
not run
B122What is the point prevalence of COPD on January 1, 2023? Report the prevalence.
Trap and rubric

Point-in-time denominator: point prevalence on a specific date counts, in the numerator, patients with a COPD diagnosis in a defined lookback (e.g. prior 3 years) who are ENROLLED on that exact date; the denominator is patients enrolled on that exact date. A one-shot query typically uses everyone with a COPD code ever / all patients, ignoring the on-date enrollment requirement for both numerator and denominator.

Grading: Must restrict both numerator and denominator to patients enrolled on 2023-01-01 (enrollment span covering that date), with the numerator having a COPD diagnosis in the prior window. Using all-time codes / all patients, ignoring on-date enrollment, is the error.

  • PASS Numerator and denominator both restricted to enrollment covering 2023-01-01
  • FAIL Uses any COPD code ever over all patients, ignoring on-date enrollment
PT1-pointprev
proportion · point-in-time-denominator
not run
B123What is the prevalence of asthma in 2022? Report the prevalence.
Trap and rubric

Prevalence-definition ambiguity: 'prevalence in 2022' can mean PERIOD prevalence (any asthma diagnosis during 2022 among those enrolled in 2022) or POINT/annual prevalence with a lookback (asthma dx in a prior window while enrolled). These give materially different counts. A defensible system states which definition it uses and applies matching enrollment; a one-shot query silently picks 'any code in 2022 / all patients' without acknowledging the choice or aligning the denominator.

Grading: Must adopt a coherent prevalence definition (period or point) and align the denominator to the enrolled population for that definition, stating the choice. Silently counting any asthma code in 2022 over all patients (mismatched denominator, undeclared definition) is the error.

  • PASS Picks period or point prevalence explicitly and aligns the enrolled denominator
  • FAIL Counts any 2022 asthma code over all patients, undeclared/mismatched
PT2-periodvspoint
proportion · interpretation-divergence
not run
B124What is the incidence of gout per 1,000 person-years over 2020-2022? Report the incidence.
Trap and rubric

Multi-span person-time: patients often have MULTIPLE enrollment spans (disenroll then re-enroll). Person-years at risk must be SUMMED across all qualifying spans within 2020-2022 (and stop at the first gout event), not taken from only the latest/longest span. A one-shot query typically uses a single span or a naive max-min date range, mis-stating the denominator person-time.

Grading: Must sum at-risk person-time across each patient's multiple enrollment spans within the window (censoring at first gout, disenrollment, or window end), and count incident (first) gout. Using a single span or last-minus-first dates for person-time is the error.

  • PASS Sums person-time across all enrollment spans, censoring at first event
  • FAIL Uses one span / naive date range for person-time
PT3-multispan
rate · person-time-multispan
not run
B125Calculate outpatient-observable person-time for 2022, excluding days each patient spent as an inpatient. Report total person-years.
Trap and rubric

Person-time with carved-out intervals: Inpatient Visit stays must be SUBTRACTED from each patient's observable interval (observation_period intersected with 2022), which requires interval difference against the union of inpatient visit_occurrence intervals, handling overlapping/adjacent stays so no inpatient day is double-subtracted. A one-shot query subtracts a raw count of inpatient rows/days or ignores inpatient time.

Grading: Must subtract the MERGED Inpatient Visit intervals (visit_start_date/visit_end_date) from each patient's observable interval (interval difference), not a raw inpatient-day count, then sum remaining person-time. Subtracting raw inpatient rows (overlap double-count) or ignoring inpatient time is the error.

  • PASS Observable interval minus merged inpatient intervals (no double-subtract)
  • FAIL Subtracts raw inpatient row/day counts, mishandling overlaps
PTINP1-excludeinpatient
summary-statistic · person-time-exclusion
not run
B126Split each patient's observable time across calendar month, calendar year, age group, and diabetes disease-status (pre- vs post-diagnosis) boundaries, and report total person-days by each combination without double-counting any day.
Trap and rubric

Multidimensional person-time splitting: each observable interval must be cut simultaneously at month-ends, year-ends, birthdays, AND the diagnosis date, so every day belongs to exactly ONE (month, year, age, disease-status) cell with no day counted twice or dropped. A one-shot query splits on at most one dimension, or assigns whole intervals to a single cell, producing overlapping/duplicated or missing person-days. Correct handling intersects all boundary sets and verifies the day total reconciles.

Grading: Must split intervals at the union of month, year, birthday, and diagnosis boundaries so each day maps to exactly one multidimensional cell (total person-days reconciles to raw observable days). Splitting on one dimension or assigning whole intervals is the error.

  • PASS Cuts at all boundary types; each day in exactly one cell; total reconciles
  • FAIL Splits on one dimension / assigns whole intervals, duplicating or dropping days
PTMD1-multidim
table · person-time-multidim
not run
B127Compute each patient's total observable person-days two ways — by expanding to one row per observable day and by summing merged date intervals — and report whether the two totals agree.
Trap and rubric

Person-time method reconciliation: the daily-expansion total and the merged-interval total must AGREE only if overlapping enrollment spans are correctly merged and interval endpoints are counted consistently (inclusive/exclusive). Discrepancies reveal double-counted overlap days or off-by-one endpoint errors. A one-shot query computes one method (often raw interval sums that double-count overlaps) and never cross-checks. Correct handling implements both and reconciles, exposing overlap/boundary handling.

Grading: Must compute person-days by day-level expansion AND by summing merged intervals, then compare — agreement requires merging overlaps and consistent endpoint counting. Reporting a single method (esp. raw interval sums that double-count overlaps) with no reconciliation is the error.

  • PASS Computes both daily-expansion and merged-interval totals and reconciles
  • FAIL One method only (raw sums double-count overlaps), no cross-check
PTREC1-reconcile
table · person-time-reconciliation
not run
B128Report total observed person-years by calendar year and age group, correctly splitting each patient's observable time when they cross a year boundary or have a birthday mid-interval. Report the grid.
Trap and rubric

Person-time splitting across year AND age boundaries: a single observable interval must be CUT at each Dec 31 and at each birthday, allocating the right fraction of person-time to each (year, age-group) cell. A one-shot query assigns a patient's whole interval to one year/age (by index date or enrollment start), mis-allocating person-time. Correct handling splits intervals at both boundary types.

Grading: Must split each observable interval at calendar-year boundaries and at birthdays, attributing the correct person-time fraction to each (year, age-group) cell. Assigning the whole interval to a single year/age bucket is the error.

  • PASS Splits intervals at year-ends and birthdays; allocates fractional person-time
  • FAIL Assigns each interval to one year/age by index/start date
PY1-yearage
table · person-time-split
not run
B129Report both the annual and the monthly counts of active heart-failure episodes for 2022, and reconcile them — an episode spanning multiple months must not be summed across months to equal the annual total. Report both plus the reconciliation.
Trap and rubric

Cross-grain reconciliation: an episode active across several months is counted in EACH of those months, so summing monthly counts overstates the (distinct-episode) annual total. The system must compute annual as distinct episodes active in the year, monthly as episodes active per month, and explain the difference (episodes spanning months). A one-shot query reports one grain, or naively sums months to 'annual', producing an inflated, inconsistent total.

Grading: Must compute the annual count as distinct episodes active in 2022 AND monthly active-episode counts, recognizing that monthly counts sum to MORE than annual because multi-month episodes recur; the totals are reconciled, not equated. Summing months to get the annual total is the error.

  • PASS Distinct-episode annual + per-month active counts, reconciled (not summed)
  • FAIL Sums monthly counts to the annual total, over-counting multi-month episodes
REC1-reconcile
table · aggregation-reconciliation
not run
B130How many acute myocardial infarction events occurred in 2022, where events for the same patient within 30 days of each other count as a single event? Report the event count.
Trap and rubric

Recurrent-event counting with a refractory window: the same clinical event generates repeat codes across days; a NEW event only counts if it is >30 days after the prior counted event for that patient. This differs from first-event-only (undercount) and from raw code counts (overcount). It also differs from episode-collapsing by adjacency — here re-qualification requires a fixed clearance gap. Correct handling walks each patient's ordered event dates, starting a new event only when >30 days have elapsed since the last counted one.

Grading: Must count recurrent events per patient where a new event requires >30 days since the previously counted event (collapsing codes within 30 days into one), allowing multiple events per patient. Counting all codes (overcount) or only the first event (undercount) is the error.

  • PASS Ordered event dates; new event only if >30 days since last counted
  • FAIL Counts every code (overcount) or only the first event per patient
RECUR1-refractory
count · recurrent-event
not run
B131How many patients qualified for a heart-failure cohort, stopped qualifying, and then re-qualified at least 180 days after they last stopped qualifying? Report the count.
Trap and rubric

Requalification after a clearance gap: membership is a time-varying interval; requalification requires the patient to LEAVE the cohort and re-enter only after >=180 days of non-qualifying time. A one-shot query treats cohort membership as a static ever/never flag and cannot express leave-then-return, or counts anyone with two qualifying spans regardless of the 180-day gap. Correct handling builds per-patient qualifying intervals and detects a re-entry preceded by a >=180-day non-qualifying gap.

Grading: Must construct time-varying cohort-membership intervals per patient and count those with a re-entry occurring >=180 days after a prior exit. A static ever-qualified flag, or ignoring the 180-day gap, is the error.

  • PASS Builds membership intervals; detects re-entry after >=180d non-qualifying gap
  • FAIL Treats membership as ever/never or ignores the clearance gap
REQUAL1-requalify
count · cohort-requalification
not run
B132For a single named maintenance medication (levothyroxine), reconstruct exposure when refills arrive before the prior supply is exhausted, carrying unused supply forward. Report the median continuous exposure duration.
Trap and rubric

Refill stockpiling / carry-forward: early refills mean supplied days accumulate ahead of consumption; a correct exposure timeline shifts each levothyroxine drug_exposure's start to the end of the prior adjusted interval (carry-forward), rather than starting at the raw drug_exposure_start_date. A one-shot query sums raw interval lengths (double-counting overlap) or uses fill-date-to-fill-date spacing.

Grading: Must build the exposure timeline by carrying forward unused supply (each exposure's coverage begins when the prior adjusted interval ends), then measure continuous exposure. Summing raw drug_exposure interval lengths over overlaps, or raw start-date spacing, is the error.

  • PASS Shifts fill coverage start by leftover supply (stockpiling-adjusted timeline)
  • FAIL Sums raw days-supply / uses raw fill dates, double-counting overlap
RX1-stockpile
summary-statistic · refill-stockpiling
not run
B133Report the distribution of the Charlson Comorbidity Index among patients hospitalized for pneumonia. Report the distribution.
Trap and rubric

Composite clinical score from many components: the Charlson index is computed per patient by mapping ~17 weighted comorbidity categories (MI, CHF, dementia, diabetes, cancer, liver disease, etc.) from diagnosis codes and summing the weights. A one-shot query cannot assemble a 17-category weighted score in a single pass — it tends to count raw diagnoses or a single condition. Correct handling builds each comorbidity flag per patient, applies the Charlson weights, sums per patient, then reports the score distribution.

Grading: Must map the ~17 Charlson comorbidity categories from diagnosis codes per patient, apply the standard category weights, sum to a per-patient score, and report the distribution. Counting raw diagnoses or omitting the weighted multi-category composition is the error.

  • PASS Builds all Charlson categories per patient, applies weights, sums, reports distribution
  • FAIL Counts raw diagnoses or a subset without the weighted 17-category composite
SC1-charlson
table · composite-score
not run
B134Is there a seasonal pattern in new influenza diagnoses? Report the diagnosis rate by calendar month, pooled across 2019-2023.
Trap and rubric

Seasonality with person-time / days-in-month normalization: raw monthly COUNTS conflate true seasonality with (a) unequal days per month (Feb vs Jul) and (b) changing enrollment/observable population by month/year. A one-shot query reports raw counts per month, implying a seasonal signal that is partly a denominator artifact. Correct handling divides monthly events by observable person-time (or population) that month, and pools rates comparably across months.

Grading: Must express each month as a RATE normalized by observable person-time/population that month (accounting for days-in-month and enrollment), not raw counts, before comparing months. Reporting raw monthly counts as the seasonal signal is the error.

  • PASS Monthly events / observable person-time (normalized for days & enrollment)
  • FAIL Raw counts per month, conflating seasonality with denominator size
SEAS1-monthnorm
table · seasonality-normalization
not run
B135Convert each patient's overlapping active chronic-condition periods into non-overlapping time segments, each labeled by the exact set of conditions active during that segment. Report the segment breakdown.
Trap and rubric

Interval flattening into labeled segments: overlapping condition intervals must be cut at every start/stop boundary into disjoint segments, each tagged with the SET of conditions simultaneously active. A one-shot query cannot decompose overlapping intervals into a segment timeline; it reports per-condition durations that overlap and double-count time. Correct handling is a sweep-line / boundary-split over interval unions.

Grading: Must split the union of overlapping condition intervals at every boundary into non-overlapping segments, each labeled by the active condition set, so total segment time has no double-counting. Reporting overlapping per-condition intervals is the error.

  • PASS Sweep-line splits overlaps into disjoint segments labeled by active condition set
  • FAIL Reports per-condition durations that overlap/double-count time
SEG1-flatten
table · interval-flatten
not run
B136Count patients who had, in order, a qualifying diagnosis, then a confirmatory test, then a treatment initiation, all within a single 90-day window (unrelated events may occur in between). Report the count.
Trap and rubric

Ordered subsequence within a window: the three events must occur in the specified ORDER and within 90 days of each other, but other unrelated events may interleave. A one-shot query checks co-occurrence of the three within 90 days without enforcing order, or requires them to be strictly consecutive records. Correct handling finds an ordered subsequence (dx < test < treatment) inside a 90-day span per patient.

Grading: Must require dx-date < test-date < treatment-date with the span from dx to treatment <=90 days, per patient, allowing unrelated intervening events. Ignoring order, or requiring strict adjacency, is the error.

  • PASS Enforces dx<test<treatment within a 90-day span, intervening events allowed
  • FAIL Checks co-occurrence without order, or requires strict adjacency
SEQ1-subsequence
count · ordered-subsequence
not run
B137How many patients completed a full 3-dose hepatitis B vaccination series, respecting the minimum required intervals between doses? Report the count.
Trap and rubric

Dose-series construction with spacing rules: a valid series requires 3 doses in order with MINIMUM intervals between them (e.g. dose2 >= 4 weeks after dose1, dose3 >= 8 weeks after dose2 and >= 16 after dose1); doses too close don't count and same-day duplicates collapse. A one-shot query counts patients with >=3 vaccine records, ignoring spacing and duplicates. Correct handling walks the ordered doses applying the interval rules.

Grading: Must validate an ordered 3-dose series meeting the minimum inter-dose intervals (collapsing same-day duplicates, rejecting too-close doses). Counting patients with >=3 vaccine records regardless of spacing is the error.

  • PASS Validates ordered doses with minimum intervals, dedups same-day
  • FAIL Counts >=3 vaccine records ignoring spacing/duplicates
SER1-doseseries
count · series-construction
not run
B138Using the code-based, lab-based, and medication-based definitions of diabetes, report how many patients meet exactly one, exactly two, or all three definitions. Report the mutually exclusive counts.
Trap and rubric

Mutually-exclusive set partitioning (no double-count): a patient meeting 2 diabetes definitions (condition_occurrence codes, measurement-based lab thresholds, and drug_exposure) must be counted once in the 'exactly two' bucket, not in each. Requires per-patient membership across the three definitions and partitioning by the COUNT of definitions met. A one-shot query reports each definition's total (overlapping) or a simple union, which double-counts.

Grading: Must compute, per patient, how many of the three diabetes definitions are met, then bucket patients by that count (exactly 1 / exactly 2 / all 3) with each patient in exactly one bucket. Reporting overlapping per-definition totals is the error.

  • PASS Counts definitions met per patient; partitions into mutually exclusive buckets
  • FAIL Reports per-definition totals or a union, double-counting
SET1-exclusive
table · set-partitioning
not run
B139Among patients with any malignancy diagnosis, how many have non-melanoma skin cancer as their ONLY malignancy? Report the count.
Trap and rubric

Patient-level set subtraction, not code-level: 'only malignancy is NMSC' means the patient has >=1 NMSC code AND ZERO codes for any other malignancy — a per-PATIENT condition over their full code set, not a per-code filter. A one-shot query filters to NMSC codes (keeping patients who also have other cancers) or subtracts at the code level, both wrong. Correct handling computes, per patient, the set of distinct malignancy types and keeps those whose set == {NMSC}.

Grading: Must identify patients whose ENTIRE set of malignancy diagnoses contains only non-melanoma skin cancer (>=1 NMSC and no other malignancy code), a patient-level set condition. Filtering to NMSC rows (ignoring patients' other cancers) or code-level subtraction is the error.

  • PASS Keeps patients whose full malignancy set is exactly {NMSC}
  • FAIL Filters to NMSC codes, keeping patients with other cancers too
SET2-onlysubset
count · set-subtraction
not run
B140How many patients have chronic kidney disease stage 3 or worse? Report the count.
Trap and rubric

Lab-based staging with repeated-measure confirmation: CKD stage 3+ is defined physiologically by TWO eGFR measurements < 60 taken at least 90 days apart (to establish chronicity), not by a single low eGFR and not by CKD stage billing codes (under-coded). A one-shot query uses stage codes, or a single eGFR<60 (which may be acute kidney injury, not chronic). Correct handling requires two qualifying eGFR values >=90 days apart per patient.

Grading: Must define CKD 3+ from >=2 eGFR measurements <60 at least 90 days apart per patient (chronicity), computing eGFR from creatinine if needed — not from CKD stage codes and not from a single low eGFR. Single-measurement or code-based definitions commit this error.

  • PASS Requires two eGFR<60 >=90 days apart per patient (chronic)
  • FAIL Uses CKD stage codes, or a single eGFR<60 (may be acute)
ST1-eGFRstage
count · lab-based-staging
not run
B141How many patients had at least three consecutive abnormal laboratory results with no normal result in between? Report the count.
Trap and rubric

Consecutive-run / streak detection: must order each patient's results by date and find a run of >=3 abnormal results uninterrupted by any normal result. A one-shot query counts patients with >=3 abnormal results total (ignoring that a normal result in between breaks the streak). Correct handling is gap-and-island/streak logic over the ordered sequence.

Grading: Must detect a maximal run of >=3 consecutive abnormal results (ordered by date) with no intervening normal result per patient. Counting >=3 abnormal results anywhere (streak-agnostic) is the error.

  • PASS Finds >=3 consecutive abnormals with no intervening normal (ordered)
  • FAIL Counts >=3 abnormal results total, ignoring intervening normals
STK1-streak
count · consecutive-run
not run
B142Report the prevalence of obesity by calendar year and sex for 2020-2022. Some patients have sex recorded inconsistently across years.
Trap and rubric

Inconsistent time-varying attribute across strata: when a patient's recorded sex differs across years, naive stratification places the same patient in different sex strata in different years (or double-counts), producing incoherent sex-specific denominators. A one-shot query strata by the per-row/per-year value without reconciling. Correct handling resolves each patient to ONE sex (e.g. most-recent or modal, stated) and applies it consistently, or explicitly reports the inconsistency count.

Grading: Must resolve each patient's sex to a single value (stated rule, e.g. modal/most-recent) applied consistently across years, or flag inconsistent patients as a separate bucket — not let the same patient flip strata year to year. Stratifying by unreconciled per-year sex is the error.

  • PASS Resolves each patient to one sex (stated rule) applied across all years
  • FAIL Uses per-year sex, letting patients flip strata / double-count
STR1-attrconsistency
table · strata-consistency
not run
B143How many patients had four or more emergency-department visits within any rolling 90-day period? Report the count.
Trap and rubric

Sliding-window maximum (not fixed calendar windows): 'any rolling 90-day period' means the window can start on ANY visit date, so a patient with 4 visits spanning day 10-95 qualifies even though no single calendar quarter contains all four. A one-shot query buckets by fixed calendar quarter/year and misses cross-boundary clusters. Correct handling checks, per patient, whether any visit-anchored 90-day window contains >=4 visits.

Grading: Must evaluate rolling 90-day windows anchored at each ED visit per patient (e.g. count visits within 90 days of each visit) and flag patients reaching >=4 in any such window. Fixed calendar-period bucketing is the error.

  • PASS Checks visit-anchored rolling 90-day windows per patient
  • FAIL Buckets visits by fixed calendar period, missing cross-boundary clusters
SW1-slidingmax
count · sliding-window-max
not run
B144Classify each patient's terminal state as death, end of observability (disenrollment), or end of available data, and report the distribution.
Trap and rubric

Terminal-state classification: a patient's follow-up ends for one of three DISTINCT reasons — death, disenrollment, or the data cut-off — and conflating them (e.g. treating disenrollment or data-end as death, or all non-deaths as 'still followed') misrepresents censoring. A one-shot query typically has no notion of why observation stops. Correct handling assigns each patient exactly one terminal reason using the earliest applicable of death date, disenrollment date, and global data-end.

Grading: Must classify each patient's follow-up end as death vs disenrollment vs data-end (mutually exclusive, earliest applicable), not conflate them. Ignoring the reason observation stops, or equating disenrollment/data-end with death, is the error.

  • PASS Assigns death / disenrollment / data-end by earliest applicable, exclusive
  • FAIL Conflates the terminal reasons or ignores why follow-up ends
TERM1-terminalstate
table · terminal-state-classification
not run
B145For each patient, identify their first qualifying diabetes diagnosis and report the count by diagnosis type; when multiple qualifying diagnoses occur on the same first date, count the patient once.
Trap and rubric

Same-day tie-break for 'first' event: when a patient has multiple qualifying diagnoses on their earliest date, 'the first' is ambiguous — a naive MIN(date) join returns MULTIPLE rows for that patient, double-counting them across diagnosis types. A one-shot query joins on the minimum date without collapsing ties, inflating per-type counts and total > patient count. Correct handling applies a deterministic tie-break (or counts the patient once) so each patient contributes exactly one first event.

Grading: Must ensure each patient contributes exactly one first event despite same-date ties (deterministic tie-break or count-once), so per-type counts sum to the distinct patient count. A MIN(date) join that returns multiple same-day rows per patient is the error.

  • PASS Resolves same-day ties so each patient has one first event
  • FAIL MIN(date) join returns multiple same-day rows, double-counting patients
TIE1-samedaytie
table · first-event-tiebreak
not run
B146Produce a transition matrix of counts of observed movements between consecutive recorded CKD stages across all patients (e.g. stage 2 to stage 3), collapsing consecutive same-stage records first. Report the matrix.
Trap and rubric

Consecutive-state transition counting: a transition is a change between a patient's TEMPORALLY ADJACENT distinct states after collapsing repeats — so stage2,stage2,stage3 yields one 2->3 transition, not two. A one-shot query cross-joins all stage pairs per patient (counting non-adjacent and self pairs) or counts every record pair, inflating the matrix. Correct handling orders states, collapses runs, and tallies adjacent (from,to) pairs.

Grading: Must collapse consecutive same-stage records, then count only temporally adjacent (from->to) stage pairs into the matrix. Cross-joining all stage pairs or counting non-collapsed adjacent duplicates is the error.

  • PASS Collapses runs; tallies temporally adjacent from->to pairs
  • FAIL Counts all stage pairs / uncollapsed duplicates, inflating the matrix
TRANS1-matrix
table · transition-matrix
not run
B147Count patients who had an abnormal cardiac stress test that was followed by a coronary angiography within 90 days, and then a revascularization procedure. Report the count.
Trap and rubric

Event-ordering dependency: the answer requires a strict temporal SEQUENCE — abnormal stress test, THEN angiography within 90 days, THEN revascularization (after the angiography). A one-shot query typically writes flat co-occurrence filters and ignores ordering, counting patients who had all three in any order. Correct handling reasons about per-patient event order and inter-event windows.

Grading: System must enforce the temporal order (stress test -> angiography within 90d -> later revascularization) using event dates per patient. Counting patients who have all three procedures without ordering/windowing is the error.

  • PASS Enforces stress->angiography(<=90d)->revascularization order per patient
  • FAIL Counts patients with all three procedures regardless of order/timing
TS1-order
count · temporal-sequence
not run
B148Count patients whose first-ever diagnosis of chronic kidney disease occurred AFTER their first diagnosis of type 2 diabetes. Report the count.
Trap and rubric

First-event ordering: the answer depends on comparing the FIRST occurrence date of two conditions per patient (CKD onset after T2D onset). A one-shot query that just requires both diagnoses present (or compares any CKD date to any T2D date) mis-answers. Correct handling computes MIN(date) per condition per patient and compares the anchors.

Grading: System must compute each patient's first (earliest) T2D date and first CKD date and count only those where first CKD > first T2D. Using any-date co-occurrence, or not anchoring on first occurrence, is the error.

  • PASS Compares MIN(CKD date) > MIN(T2D date) per patient
  • FAIL Counts patients with both diagnoses without first-occurrence ordering
TS2-order
count · temporal-sequence
not run
B149Count patients whose first opioid prescription came AFTER their first documented chronic pain diagnosis (not before). Report the count.
Trap and rubric

First-event ordering across two domains: requires comparing MIN(chronic-pain condition_occurrence date) to MIN(opioid drug_exposure date) per patient and keeping only opioid-after-pain. A one-shot query that merely requires both present, or compares arbitrary dates, mis-answers.

Grading: Must compute first pain-diagnosis date (condition_occurrence) and first opioid drug_exposure date per patient and count only those with first opioid > first pain. Co-occurrence or non-first-event date comparison is the error.

  • PASS Compares MIN(opioid date) > MIN(pain dx date) per patient
  • FAIL Counts patients with both, ignoring first-event ordering
TS3-order
count · temporal-sequence
not run
B150Count patients diagnosed with lung cancer whose diagnosis was preceded by a chest CT within the prior 90 days (a workup-then-diagnosis pattern). Report the count.
Trap and rubric

Directional pre-event window (ordering): the CT must come BEFORE the cancer diagnosis, within 90 days prior. A one-shot query commonly checks 'has lung cancer AND has chest CT' or uses a symmetric window, capturing post-diagnosis surveillance CTs too. Correct handling enforces CT-before-diagnosis within the prior-90-day window.

Grading: Must require a chest CT dated within the 90 days BEFORE the lung-cancer diagnosis, per patient (directional). Counting any CT (including after diagnosis) or ignoring the window is the error.

  • PASS Requires chest CT in the 90 days before the cancer diagnosis
  • FAIL Counts any lung-cancer patient with any chest CT, ignoring direction/window
TS4-order
count · temporal-sequence
not run
B151Among patients newly diagnosed with heart failure, what is the median time to first hospitalization? Report the median in days.
Trap and rubric

Time-to-event with right-censoring: patients who disenroll or reach data end WITHOUT being hospitalized are censored, not absent. Taking the median only over patients who were hospitalized (ignoring censored follow-up) badly underestimates the median time — and the true median may be un-reached (>50% never hospitalized). A one-shot query computes the mean/median of observed event times among those with the event. Correct handling accounts for censored follow-up (Kaplan-Meier median or explicit at-risk reasoning), or states the median is not reached.

Grading: Must incorporate censored follow-up (patients without the event contribute at-risk time; median from a survival/at-risk estimate, or reported as not-reached if <50% have the event). Taking the median of event times only among those hospitalized is the error.

  • PASS Uses survival/at-risk logic incl. censored patients (or reports not-reached)
  • FAIL Median over hospitalized patients only, ignoring censoring
TTE1-censored
summary-statistic · time-to-event-censoring
not run
B152For heart failure in 2022, report the count at five levels: distinct patients, distinct hospitalization episodes, distinct encounters, distinct records, and records. Report all five.
Trap and rubric

Unit-of-analysis integrity: the same clinical reality yields very different numbers at patient / hospitalization-episode / visit / condition-record grain, and conflating them is a classic RWE error. The question forces each grain to be computed correctly (distinct persons, collapsed inpatient episodes, distinct visit_occurrence, raw condition_occurrence rows). A one-shot query returns one COUNT(*) at whatever grain the table happens to be.

Grading: Must produce distinct counts at the patient, hospitalization-episode (collapsed Inpatient Visits), visit, and condition-record grains — each at its correct level of aggregation. A single COUNT at one grain is the error.

  • PASS Computes distinct-patient, episode, visit, and condition-record counts separately
  • FAIL Returns one COUNT(*) without distinguishing units of analysis
UOA1-grid
table · unit-of-analysis
not run
B153How many patients had a follow-up hepatitis B vaccination within one year of their first dose? Report the count under each interpretation of 'within one year': within 365 days, within 366 days on a leap year, and on or before the same calendar date next year.
Trap and rubric

Competing definitions of 'within one year': '365 days', 'the same calendar date next year' (which is 366 days across a leap year), and 'by the first anniversary' are DIFFERENT boundaries yielding different counts for events near the edge. A one-shot query silently equates 'one year' with a single arithmetic (usually 365 days) and hides the divergence. Correct handling computes each interpretation and surfaces that the answer depends on the definition.

Grading: Must compute the count under each distinct 'within one year' interpretation (365-day, calendar-anniversary incl. leap-year effect) and surface the divergence, not silently pick one. Collapsing all senses into a single 365-day rule is the error.

  • PASS Computes each 'within one year' definition and reports the divergence
  • FAIL Silently uses one interpretation (365 days), hiding the difference
WITHIN1-oneyear
table · interpretation-divergence
not run

Tasks that create objects in Linkr, which the user then reviews and edits in the interface. Database: MIMIC-IV demo in OMOP CDM. Expected results are produced by hand, independently of the runs. Three runs per task. Values in brackets are not fixed yet.

Prompt copies the instruction to paste into a new conversation of an MCP client connected to Linkr.

#TaskExpected outputSuccessResult
C1Cohort: adults with an ICU stay of 48 hours or more
Patient listSame listnot run
C2Cohort: vancomycin given in the ICU
Patient listSame listnot run
C3Cohort: lactate above 4 mmol/L within 24 hours of ICU admission
Patient listSame listnot run
C4Edit C1 to exclude deaths within 24 hours, report the attrition
Patient list, attritionBoth identicalnot run
C5Import an OHDSI Phenotype Library definition and run it
Patient listSame listnot run
C6Map 20 laboratory source codes
OHDSI mapping[threshold] correctnot run
C7Map 20 drug source codes
OHDSI mapping[threshold] correctnot run
C8Dataset from C1: age, sex, length of stay, first-day maximum lactate, documented
TableSame valuesnot run
C9R script: descriptive table of C1
TableSame valuesnot run
C10Python script: distribution of the first lactate, saved figure
Quartiles, fileSame quartiles, file presentnot run
C11Dashboard for C1: patient count, age histogram, sex distribution
3 widgets3 of 3not run
#HarnessModelAnswerVerdictToolsTokens inBatch

Click a row to see the prompt, the tool calls, the model's last reply and the verdict.