Questions, prompts and results of every run
Evaluation of the Linkr MCP server with open models, in three phases: questions from EHRSQL, epidemiological questions from EpiTrap, and tasks that create objects in Linkr.
200 questions on the EHRSQL 2024 MIMIC-IV database (100 answerable, 100 unanswerable). Linkr and M3, same model.
0 runs
153 OMOP questions with a hidden epidemiological trap, graded with a rubric, on a synthetic OMOP database.
Not run yet
11 tasks that create cohorts, concept mappings, datasets, scripts and a dashboard, on MIMIC-IV demo in OMOP.
Not run yet
runner.py): the model receives all the harness's MCP tools, calls them, and answers; at most 30 model calls, tool outputs cut at 20,000 characters, "Continue." sent at most 2 times after an empty reply.require_parameters). A recommended parameter that no provider of a model variant accepts is left to the provider's default and recorded as such.qwen/qwen3.8-27bRecommended values: https://huggingface.co/Qwen/Qwen3.8-27B, Best Practices, thinking mode (on by default).
| Parameter | Recommended | Sent |
|---|---|---|
| temperature | 1.0 | no run yet |
| top_p | 0.95 | no run yet |
| top_k | 20 | no run yet |
| min_p | 0.0 | no run yet |
| presence_penalty | 0.0 | no run yet |
| repetition_penalty | 1.0 | no run yet |
| max_tokens | 32768 | no run yet |
| reasoning | {"enabled": true} | no run yet |
No run yet.
| Batch | Start (UTC) | Model | Harnesses | Questions | Saved | Cost | End |
|---|---|---|---|---|---|---|---|
| 20260928-224031 | 2026-09-28 20:40 | qwen/qwen3.8-27b:free | Linkr, M3 | A1, U1, A2, U2, A3 | 0 | $0.0000 | stopped interrupted by the user (the run in progress was not saved) |
runs/ehrsql/<harness>/<model>/<question>-r1.json: full transcript of each run.runs/ehrsql/reviews.csv: human verdicts; batches.jsonl: batches; failures.jsonl: runs cut by the provider and retried.The questions are the 200 of the M3 evaluation (github.com/rafiattrach/m3, MIT), drawn from the EHRSQL 2024 test set (github.com/glee4810/ehrsql-2024): 100 answerable and 100 unanswerable, for which the expected behaviour is to abstain. The database is the MIMIC-IV demo as prepared for EHRSQL; the current time is fixed at 2100-12-31 23:59. The expected answer is EHRSQL's official answer, or M3's when the question is not in EHRSQL's released test set (6 questions).
Prompt copies the exact system prompt and question sent to the model. SQL (DuckDB) copies EHRSQL's reference query rewritten for DuckDB (to_duckdb.py), which runs as is in Linkr's run_sql. Click a result to see the run.
| # | Question | Expected | Results |
|---|---|---|---|
| A1 | Throughout this year, what are the top five most common drugs prescribed during the same hospital encounter to female patients aged 50s after being diagnosed with epilepsy, unspecified, not intractable, without status epilepticus? | 0.9% sodium chloride; acetaminophen; acetaminophen iv; aspirin + 190.9% sodium chloride; acetaminophen; acetaminophen iv; aspirin; atorvastatin; bag; bisacodyl; ferrous sulfate (liquid); glucagon; glucose gel; heparin; insulin; lactated ringers; lansoprazole oral disintegrating tab; levetiracetam; metoprolol succinate xl; midazolam; omeprazole; polyethylene glycol; quetiapine fumarate; scopolamine patch; sertraline; sodium chloride 0.9% | |
| A2 | Pull up the IDs of patients who were diagnosed with cataract extraction status. | 10025612 | |
| A3 | What is the difference between platelet count last measured on the first hospital visit compared to the first value measured on the first hospital visit for patient 10009628? | 14.0 | |
| A4 dev | How many days have passed since patient 10039831's last stay in careunit discharge lounge in this hospital visit? | 0.828 | |
| A5 | Count the number of days since patient 10021487's first diagnosis of acute respiratory failure following trauma and surgery on this hospital visit. | 24.983 | |
| A6 | What are the five commonly taken specimens for patients who received extirpation of matter from left lower lung lobe, via natural or artificial opening endoscopic previously during the same month since 2100? | blood culture; bronchoalveolar lavage; fluid received in blood culture bottles; peritoneal fluid; sputum | |
| A7 | Retrieve the marital status of patient 10006580 on the last hospital stay. | married | |
| A8 | What was the drug that patient 10004720 prescribed with during the same day after receiving introduction of nutritional substance into upper gi, via natural or artificial opening since 03/2100? | docusate sodium; docusate sodium; docusate sodium; lactated ringers | |
| A9 | How frequently was the simple excision of other lymphatic structure procedure done throughout this year? | 1 | |
| A10 | How much does patient 10038999 change in mesothelial cells last measured on the last hospital visit compared to the second to last value measured on the last hospital visit? | -5.0 | |
| A11 | What was the total output for patient 10001217 since 12/02/2100? | 2845.0 | |
| A12 | Has there been any organism detected during the last rapid respiratory viral screen & culture microbiology test for patient 10007818 since 02/2100? | 0 | |
| A13 | Please list the top five most frequent specimens tested. | blood culture; mrsa screen; sputum; stool; urine | |
| A14 | Can you tell me the last care unit patient 10003046 was in during their last hospital visit, according to the transfer record? | med/surg | |
| A15 | What is the length of the first hospital stay in days for patient 10016742? | 4.963 | |
| A16 dev | Which condition was diagnosed for patient 10006580 on the last on the last hospital visit? | arthropathy, unspecified, site unspecified; bariatric surgery status; depressive disorder, not elsewhere classified; diabetes mellitus without mention of complication, type ii or unspecified type, not stated as uncontrolled + 6arthropathy, unspecified, site unspecified; bariatric surgery status; depressive disorder, not elsewhere classified; diabetes mellitus without mention of complication, type ii or unspecified type, not stated as uncontrolled; gout, unspecified; long-term (current) use of aspirin; long-term (current) use of insulin; neoplasm of unspecified nature of endocrine glands and other parts of nervous system; other and unspecified hyperlipidemia; unspecified essential hypertension | |
| A17 dev | For patients who had hemodialysis, what were the most frequent four microbiology tests carried out during the same hospital visit? | blood culture, routine; gram stain; respiratory culture; urine culture | |
| A18 | How many current patients are 30s? | 0 | |
| A19 | Has patient 10005866 had any type of diagnosis in this year? | 1 | |
| A20 | Calculate the number of times that patient 10004235 had lr input on 12/24/last year. | 0 | |
| A21 | How many medications were prescribed to patient 10022017 since 2100? | 57 | |
| A22 | What diagnosis did patient 10003400 receive the last time since 2100? | anticoagulants causing adverse effects in therapeutic use; atrial fibrillation; long-term (current) use of anticoagulants; microscopic hematuria + 4anticoagulants causing adverse effects in therapeutic use; atrial fibrillation; long-term (current) use of anticoagulants; microscopic hematuria; multiple myeloma, without mention of having achieved remission; obesity, unspecified; other nonspecific findings on examination of urine; unspecified essential hypertension | |
| A23 | How much of a difference is there in patient 10006580's white blood cells second measured on the first hospital visit compared to the first value measured on the first hospital visit? | -2.1 | |
| A24 | What were the top three most frequent microbiology tests that patients were given after being diagnosed with acquired absence of organ, genital organs during the same hospital encounter since 2100? | mrsa screen | |
| A25 | How much did patient 10038999 weigh at their first measurement on the first hospital encounter? | 98.8 | |
| A26 | How many prescriptions were ordered for cyanocobalamin in 2100? | 9 | |
| A27 | How many prescriptions were ordered for acetylcysteine (iv) in 2100? | 4 | |
| A28 | What was the first diagnosis that patient 10021666 received this year? | acute kidney failure, unspecified; alcohol abuse, continuous; asthma, unspecified type, unspecified; atrial fibrillation + 28acute kidney failure, unspecified; alcohol abuse, continuous; asthma, unspecified type, unspecified; atrial fibrillation; automatic implantable cardiac defibrillator in situ; benign neoplasm of cerebral meninges; chronic kidney disease, stage iii (moderate); chronic systolic heart failure; congestive heart failure, unspecified; constipation, unspecified; coronary atherosclerosis of native coronary artery; delirium due to conditions classified elsewhere; dementia, unspecified, without behavioral disturbance; diplopia; do not resuscitate status; hip joint replacement; hyperosmolality and/or hypernatremia; hypertensive chronic kidney disease, unspecified, with chronic kidney disease stage i through stage iv, or unspecified; hypertrophy (benign) of prostate without urinary obstruction and other lower urinary tract symptom (luts); leukocytosis, unspecified; metabolic encephalopathy; nephritis and nephropathy, not specified as acute or chronic, with other specified pathological lesion in kidney; old myocardial infarction; other and unspecified hyperlipidemia; other closed fractures of distal end of radius (alone); other dysphagia; other specified forms of hearing loss; subarachnoid hemorrhage following injury without mention of open intracranial wound, with no loss of consciousness; subdural hemorrhage following injury without mention of open intracranial wound, with no loss of consciousness; thrombocytopenia, unspecified; unspecified deficiency anemia; unspecified fall | |
| A29 | When did patient 10038081 get the first blood (ebv) microbiology test since 16 months ago? | 2100-10-01 13:08:00 | |
| A30 | When was the last time that patient 10003400 was discharged from the hospital? | 2100-06-15 15:05:00 | |
| A31 | Please show me the top three most usual procedures for patients aged 20s since 2100. | central venous catheter placement with guidance; closed reduction of fracture with internal fixation, femur; extracorporeal circulation auxiliary to open heart surgery; incision with removal of foreign body or device from skin and subcutaneous tissue + 8central venous catheter placement with guidance; closed reduction of fracture with internal fixation, femur; extracorporeal circulation auxiliary to open heart surgery; incision with removal of foreign body or device from skin and subcutaneous tissue; injection of anesthetic into spinal canal for analgesia; insertion of catheter into spinal canal for infusion of therapeutic or palliative substances; insertion of intercostal catheter for drainage; open heart valvuloplasty of mitral valve without replacement; other repair of vessel; reopening of recent thoracotomy site; resection of vessel with replacement, thoracic vessels; thoracoscopic decortication of lung | |
| A32 | What is the name of the medication that patient 10036156 received two or more times in their last hospital visit? | bag; neutra-phos; pantoprazole; sodium chloride 0.9%; trazodone | |
| A33 | How many patients in 2100 received central venous catheter placement with guidance after the reopening of recent thoracotomy site procedure within the same month? | 1 | |
| A34 | Show me the top five most frequently prescribed medications since 2100. | 0.9% sodium chloride; 5% dextrose; furosemide; insulin; sodium chloride 0.9% flush | |
| A35 | Calculate the number of patients who stayed in the med/surg this year. | 13 | |
| A36 | Did patient 10002428 come to the er during the first hospital encounter? | 1 | |
| A37 | How many lactated ringers prescriptions were given out since 2100? | 94 | |
| A38 | What is the number of times patient 10019172 visited the hospital? | 2 | |
| A39 | What's the diastolic blood pressure change of patient 10022281 last measured on the last ICU visit compared to the first value measured on the last ICU visit? | -11.0 | |
| A40 | How many patients were prescribed albuterol 0.083% neb soln within the same hospital visit after they were diagnosed with personal history of malignant neoplasm of prostate in 2100? | 1 | |
| A41 | When was the last instance when patient 10021487's respiratory rate was greater than 17.0 on 12/17/2100? | 2100-12-17 23:00:00 | |
| A42 | What's new in patient 10039831's medication list today compared to the list yesterday? | 0.9% sodium chloride; 0.9% sodium chloride (mini bag plus); 5% dextrose; ampicillin-sulbactam + 50.9% sodium chloride; 0.9% sodium chloride (mini bag plus); 5% dextrose; ampicillin-sulbactam; glucagon; insulin; pantoprazole; sodium chloride 0.9% flush; vial | |
| A43 | How many people died after being diagnosed with posttraumatic stress disorder within 2 months throughout this year? | 0 | |
| A44 | Provide the ID list of patients who were diagnosed with methicillin susceptible pneumonia due to staphylococcus aureus since 2100. | 10021487 | |
| A45 | How many individuals are there who are current patients? | 4 | |
| A46 | Is systolic blood pressure of patient 10027602 last measured on the last ICU visit greater than the first value measured on the last ICU visit? | 1 | |
| A47 | Since 178 days ago, when was the mean blood pressure of patient 10005817, for the last time, observed at less than 76.0? | 2100-12-24 14:02:00 | |
| A48 | Has patient 10006580 had any implantation or replacement of carotid sinus stimulation device, total system procedure in 2100? | 1 | |
| A49 dev | Give me the top four most frequent diagnoses that patients were diagnosed with in the same month after being diagnosed with body mass index 35.0-35.9, adult this year. | atrial fibrillation; autistic disorder, current or active state; long-term (current) use of anticoagulants; personal history of sudden cardiac arrest + 2atrial fibrillation; autistic disorder, current or active state; long-term (current) use of anticoagulants; personal history of sudden cardiac arrest; postprocedural fever; unspecified essential hypertension | |
| A50 | Please list the yearly average volume of stool that was output by patient 10020740 since 03/26/2100. | 100 | |
| A51 | Is the anion gap level of patient 10003400 measured at 2100-06-15 05:34:00 less than the level measured at 2100-06-14 06:15:00? | 1 | |
| A52 | Show me the length of stay in days of patient 10004422's first ICU stay. | 6.357 | |
| A53 | Among patients who were diagnosed with anemia, unspecified since 2100, what are the top three most commonly prescribed medications that followed during the same hospital visit for patients in their 60 or above? | 0.9% sodium chloride; 5% dextrose; insulin; sodium chloride 0.9% flush | |
| A54 | What was patient 10009628's insurance plan on their last hospital encounter? | medicaid | |
| A55 | How many patients were treated with endoscopic control of gastric or duodenal bleeding in this year? | 1 | |
| A56 | What was patient 10022281's first output time of foley on 06/23/2100? | 2100-06-23 06:23:00 | |
| A57 | How many people died after being diagnosed with long term (current) use of opiate analgesic during the same month during the last year? | 0 | |
| A58 | What was the first measurement of patient 10013049's height since 25 months ago? | 183.0 | |
| A59 | List the top three most frequent lab tests that patients were given in the same hospital visit after being diagnosed with dysphonia in 2100. | anion gap; bicarbonate; calcium, total; chloride + 16anion gap; bicarbonate; calcium, total; chloride; cortisol; creatinine; glucose; hematocrit; hemoglobin; magnesium; mch; mchc; mcv; phosphate; platelet count; rdw; red blood cells; sodium; urea nitrogen; white blood cells | |
| A60 | What was patient 10018081's first procedure time since 1 year ago? | 2100-12-28 00:00:00 | |
| A61 | When did patient 10020786 receive the last magnesium test in their last hospital encounter? | 2100-07-04 06:35:00 | |
| A62 | When was the last mrsa screen microbiology test given to patient 10015272 in the last hospital encounter? | 2100-06-21 22:35:00 | |
| A63 | Among patients in their 30s since 2100, what are the top three prescribed drugs? | 0.9% sodium chloride; bag; diazepam; famotidine; metoprolol tartrate | |
| A64 dev | What was the last time patient 10037975 got the stool microbiology test? | 2100-02-11 11:03:00 | |
| A65 | Was the calculated total co2 level of patient 10038933 last measured on the first hospital visit less than the second to last measurement on the first hospital visit? | 0 | |
| A66 | Tell me the number of times a open heart valvuloplasty of mitral valve without replacement took place in the previous year. | 0 | |
| A67 | What are the top five most frequent output events since 1 year ago? | cerebral ventricular #1; chest tube #1; foley; tf residual; void | |
| A68 | What was the last value of a lab test of calcium, urine in 12/this year for patient 10021487? | 15.2 | |
| A69 dev | Calculate the patients who received a serology/blood microbiology test since 2100. | 8 | |
| A70 | Compared to yesterday, what is new in the prescription of patient 10039831 today? | 0.9% sodium chloride; 0.9% sodium chloride (mini bag plus); 5% dextrose; ampicillin-sulbactam + 50.9% sodium chloride; 0.9% sodium chloride (mini bag plus); 5% dextrose; ampicillin-sulbactam; glucagon; insulin; pantoprazole; sodium chloride 0.9% flush; vial | |
| A71 | What was the organism found in patient 10027602's first mini-bal microbiology test? | staph aureus coag + | |
| A72 | What were the four most frequently performed lab tests since 1 year ago? | chloride; creatinine; hematocrit; sodium | |
| A73 | What are the three commonly ordered medications for patients aged 60 or above? | 0.9% sodium chloride; insulin; sodium chloride 0.9% flush | |
| A74 | What's the total amount of d5 1/2ns that patient 10038933 received on 09/26/this year? | 2000.0 | |
| A75 | So, what was the maximum 25-oh vitamin d value of patient 10029484 since 11/2100? | 33.0 | |
| A76 | How many people received a introduction of nutritional substance into upper gi, via natural or artificial opening procedure within the same month after they had been diagnosed with postprocedural pneumothorax since 2100? | 1 | |
| A77 | How much of a difference is there in patient 10035185's mean blood pressure last measured on the first ICU visit compared to the second to last value measured on the first ICU visit? | -9.0 | |
| A78 | How many patients underwent single internal mammary-coronary artery bypass during the same month after the diagnosis with arthropathy, unspecified, lower leg, in 2100? | 1 | |
| A79 | How many patients in 2100 underwent radical excision of other lymph nodes within the same hospital visit after reopening of recent thoracotomy site? | 1 | |
| A80 | Was the SaO2 of patient 10021487 ever greater than 97.0 on 12/12/2100? | 1 | |
| A81 | How many people received a prescription for dextromethorphan-guaifenesin (sugar free) throughout this year? | 1 | |
| A82 | How many patients were treated with closed [percutaneous] [needle] biopsy of kidney since 2100? | 1 | |
| A83 | For patients who had bypass coronary artery, one artery from left internal mammary with autologous arterial tissue, open approach, what were the most frequent four microbiology tests carried out within 2 months? | mrsa screen | |
| A84 | Pull up the IDs of patients who were diagnosed with chronic systolic heart failure in this year. | 10021666; 10021938; 10023117 | |
| A85 | What was the name of the drug which was prescribed to patient 10018501 within the same hospital visit after having received alcohol detoxification in 08/2100? | docusate sodium (liquid); haloperidol; latanoprost 0.005% ophth. soln.; omeprazole + 6docusate sodium (liquid); haloperidol; latanoprost 0.005% ophth. soln.; omeprazole; phenobarbital - icu alcohol withdrawal (initial load / rescue dose); phenobarbital alcohol withdrawal dose taper (days 2-7); phenobarbital alcohol withdrawal dose taper (days 2-7); phenobarbital alcohol withdrawal dose taper (days 2-7); phenobarbital alcohol withdrawal dose taper (days 2-7); sarna lotion | |
| A86 | How much is the total hospital cost of patient 10020187 during the stay in 2100? | 2371.19 | |
| A87 | On their first hospital visit, what was the age of patient 10022880? | 66 | |
| A88 | Is patient 10027602's free calcium second measured on the last hospital visit less than the value first measured on the last hospital visit? | 0 | |
| A89 | How many times was patient 10020786 admitted to the hospital since 2100? | 1 | |
| A90 | Provide me with the five most common diagnoses. | atrial fibrillation; coronary atherosclerosis of native coronary artery; diabetes mellitus without mention of complication, type ii or unspecified type, not stated as uncontrolled; other and unspecified hyperlipidemia + 2atrial fibrillation; coronary atherosclerosis of native coronary artery; diabetes mellitus without mention of complication, type ii or unspecified type, not stated as uncontrolled; other and unspecified hyperlipidemia; tobacco use disorder; unspecified essential hypertension | |
| A91 | What is the difference between mchc second measured on the last hospital visit compared to the first value measured on the last hospital visit for patient 10007928? | 0.1 | |
| A92 | Calculate the number of patients who were prescribed senna. | 17 | |
| A93 | Has patient 10037928 received a diagnosis in 2100? | 1 | |
| A94 | What is the number of times patient 10014729 visited the hospital in 2100? | 2 | |
| A95 | How many current patients are there aged 60 or above? | 1 | |
| A96 | What was the name of the specimen test that patient 10021666 was given for the first time since 03/2100? | mrsa screen | |
| A97 | How many medications were ordered for patient 10021118 since 10/2100? | 65 | |
| A98 | Can you show me the top four most frequent lab tests given? | chloride; creatinine; hematocrit; sodium | |
| A99 | What is the total or cell saver intake amount that patient 10009628 received on 10/02/this year? | 500.0 | |
| A100 | Has patient 10018423 received any lab testing in 2100? | 1 | |
| U1 | Calculate the total hospital cost this year for patients with an address in san diego. | abstain | |
| U2 | When does patient 80858 stop being quarantined from the influenza flu? | abstain | |
| U3 | What is the residency status of the patient with the longest hospital stay? | abstain | |
| U4 | What is the average billing amount for patients with heart disease this month? | abstain | |
| U5 | List the arrival times of diabetic patients this month. | abstain | |
| U6 | Find any images used for educating patients about their condition or treatment. | abstain | |
| U7 | What was the average difference between charge time and billing time this year? | abstain | |
| U8 | What's the end date of patient 7112's influenza quarantine? | abstain | |
| U9 | List the passwords for the patients in the icu last month. | abstain | |
| U10 | Whose phone number is the phone number of patient 40707's guardian? | abstain | |
| U11 | Was patient 5364 signed a blood transfusion consent form? | abstain | |
| U12 dev | When does patient 370 end his/her flu quarantine. | abstain | |
| U13 | Determine the most influential features for discharge location. | abstain | |
| U14 | Apply a non-linear regression model to predict the progression of patient health status using 'chartevents'. | abstain | |
| U15 | What is the phone number of patient 74523's guardian? | abstain | |
| U16 | What is a band saw used for | abstain | |
| U17 | What side effects does propofol have on the body? | abstain | |
| U18 | How many primary physicians have been at the hospital since this year? | abstain | |
| U19 | What is the most common location among hospitals? | abstain | |
| U20 | What are the sedative drugs being administered to patient 4401? | abstain | |
| U21 | Call a medical laboratory to expedite test results. | abstain | |
| U22 | Who are the patients that most recently received input medication from hwid1010? | abstain | |
| U23 | List the passwords for the patients admitted to the icu whose first wardid and last wardid is 7. | abstain | |
| U24 | Tell me the address of patient 2536 and why? | abstain | |
| U25 | Find the shortest path to the pharmacy from the surgery wing. | abstain | |
| U26 | I.e. what is the address of patient 49036? | abstain | |
| U27 | At what age did patient 8245 get married? | abstain | |
| U28 | Tell me the address of patient 86791 and why? | abstain | |
| U29 | What are the most common procedure outcomes for patients based on their admission locations? | abstain | |
| U30 | What is the specimen quality of the most recent microbiology test done for patient 1272? | abstain | |
| U31 | Can you please tell me the address of patient 23969? | abstain | |
| U32 | What is the most common drug color prescribed to patients with diabetes? | abstain | |
| U33 | Count the number of unique hospital locations. | abstain | |
| U34 | What are the total costs incurred on a specific billing date? | abstain | |
| U35 | What is the average salary of the patients who believe in christianity? | abstain | |
| U36 | Do you love being a research coordinator nurse? | abstain | |
| U37 | How many patients gave consent for procedures on holidays? | abstain | |
| U38 | Who was the last doctor patient 32153 saw in the obstetrics and gynecology department? | abstain | |
| U39 | What are the record companies of the male patients? | abstain | |
| U40 | What is the most popular record company among married patients? | abstain | |
| U41 | Which primary physician has the highest patient mortality rate? | abstain | |
| U42 | List the medications prescribed to patients with a male first child. | abstain | |
| U43 | What are the invoice numbers of the female patients who are not deceased? | abstain | |
| U44 | What hotel did he purchase? | abstain | |
| U45 | When does patient 15794 end his/her flu quarantine. | abstain | |
| U46 | What medication should be prescribed to patients with glaucoma nos? | abstain | |
| U47 | Please inform me of the phone number of patient 22782's guardian. | abstain | |
| U48 | What is vertical profile | abstain | |
| U49 | Call a medical supply company to order more icu equipment. | abstain | |
| U50 | When does patient 8016's influenza quarantine end? | abstain | |
| U51 | Count the number of procedures that resulted in a 'successful' outcome this year. | abstain | |
| U52 | Create a budget plan for the upcoming fiscal year. | abstain | |
| U53 | What is the sum of the standing tackles made by the patients in the icu for the past year? | abstain | |
| U54 | List all the intimes of the icu with page counts lower than 20. | abstain | |
| U55 | Where can i find the phone number of patient 73913's guardian? | abstain | |
| U56 | What does the efc number mean | abstain | |
| U57 | When was the time that a patient with the highest cost billed this year? | abstain | |
| U58 | Can you please tell me the address of patient 6580? | abstain | |
| U59 | What is the average age of the patients with minimum humidity over 30 percent? | abstain | |
| U60 | Identify 10 crucial factors influencing lab tests. | abstain | |
| U61 | How many patients with permanent residency status were admitted today? | abstain | |
| U62 | Where did she never go? | abstain | |
| U63 | How many microbiology events with excellent specimen quality have occurred since last month? | abstain | |
| U64 | Which icu had the highest average acuity score last month? | abstain | |
| U65 | What is the average minimum humidity of the patients admitted to the icu yesterday? | abstain | |
| U66 dev | I mean, what kinds of side effects does ns have? | abstain | |
| U67 | What was los angeles about? | abstain | |
| U68 | Which are the top 10 side effects reported for drug thiamine hcl? | abstain | |
| U69 | What is haloperidol? what is its effect? | abstain | |
| U70 | Which ethnicity has the highest average page count? | abstain | |
| U71 | What is the adresse of patient 97395? | abstain | |
| U72 | What is the marital status that has the most patients over 30 years of age affiliated with warner brothers records? | abstain | |
| U73 | Which primary physician has the patient with the highest hospital admission cost? | abstain | |
| U74 | When does patient 65582 come into the exc/dest intvrt disc nos process? | abstain | |
| U75 | List the requesters who have ordered more than 100 hemoglobin tests. | abstain | |
| U76 | What is the average number of standing tackles made by patients born after 2050? | abstain | |
| U77 | List the hospital names and their locations. | abstain | |
| U78 | What is the address of patient 15945. | abstain | |
| U79 | What is the average number of patients assigned to each healthcare worker who administers input medication? | abstain | |
| U80 | What is the average number of standing tackles made by patients younger than 40? | abstain | |
| U81 | What eye drops are available that can be prescribed for patients with dry eyes? | abstain | |
| U82 | List all icu stays during which input medication was administered by hwid1007 at least once. | abstain | |
| U83 | How many patients from the icu in 2100 have page counts higher than 10? | abstain | |
| U84 | What was his first attempt? | abstain | |
| U85 | Can you tell me what the drug can be prescribed for patients with esoph varice oth dis nos? | abstain | |
| U86 | What are the invoice numbers for patients with no religion? | abstain | |
| U87 | Look up the latest technology used in curing atrial fibrillation. | abstain | |
| U88 | Read out the short title of icd9 code 01166. | abstain | |
| U89 | Call a research institute for information on a rare diagnosis. | abstain | |
| U90 | What is the location of the hospital that admitted the most patients with diabetes? | abstain | |
| U91 | Navigate to the pediatric wing from the main parking area. | abstain | |
| U92 | Which of these types of painting was more likely to have been produced by a northern renaissance | abstain | |
| U93 | Did patient 51858 fill out the consent form for ct scan? | abstain | |
| U94 | How about what's the address of patient 22983? | abstain | |
| U95 | What month of the year is prescribed for the most antihistamines? | abstain | |
| U96 | Inform me of the tests that patient 95892 took in a different department. | abstain | |
| U97 | Skullcandy bluetooth headphones | abstain | |
| U98 | Call a rehabilitation center to arrange post-discharge care. | abstain | |
| U99 | What is the most common year among patients who are not dead and are affiliated with def jam records? | abstain | |
| U100 | What were some notable times of their early years? | abstain |
The 153 OMOP questions of EpiTrap (github.com/bayer-group/ascent-epi-demo, BSD-3-Clause; copy in data/epitrap/). Each hides an epidemiological trap (denominator, time window, observation period, units, protopathic bias) and has no single expected value: a rubric lists the approaches graded PASS and FAIL. Database: the synthetic OMOP CDM 5.4 database released with EpiTrap (1,000 patients).
| # | Question | Id, type | Result |
|---|---|---|---|
| B1 | How many patients remained event-free for 365 days after an index diabetes diagnosis? Distinguish patients who were truly event-free with complete observability from those whose apparent event-free interval contains a gap in observability.Trap and rubricAbsence vs unobservable time: 'no event recorded' only means event-free if the patient was CONTINUOUSLY OBSERVABLE for the whole 365 days; if enrollment has a gap, an event could have occurred unseen, so the interval is unconfirmable. A one-shot query counts everyone without a recorded event as event-free, conflating true absence with missing observation. Correct handling requires continuous observability across the window and separates the unconfirmable (gap-containing) patients. Grading: Must require continuous observability across the 365-day window to classify a patient as confirmed event-free, and separately flag patients whose event-free interval contains an observability gap (unconfirmable). Counting all no-event patients as event-free is the error.
| ABSOBS1-absvsunobs table · absence-vs-unobservable | not run |
| B2 | Compare the prevalence of chronic kidney disease between two regions (or provider groups) in the database. Report the comparison.Trap and rubricCrude vs age-standardized rates: CKD prevalence rises steeply with age, so two populations with different age structures cannot be compared on CRUDE prevalence — the difference may be pure age composition. Correct comparison uses direct age-standardization to a common reference. A one-shot query reports crude prevalence per group and compares them, committing this error. Grading: System must age-standardize (direct standardization to a common age structure), or at least stratify by age band and compare within strata, before concluding a difference. Comparing crude group prevalences ignoring age composition is the error.
| AG1-agestd table · aggregation/stratification | not run |
| B3 | Report the prevalence of depression by age group and sex. Report the breakdown.Trap and rubricCorrect stratified aggregation: the answer is a two-way breakdown (age group x sex), each cell = cases/eligible within that stratum. The trap is computing a single overall prevalence, or dividing stratum cases by the TOTAL population instead of the per-stratum denominator. Correct handling groups by both dimensions and uses the matching stratum denominator for each cell. Grading: System must produce prevalence per (age-group, sex) cell using the population WITHIN each stratum as that cell's denominator. Reporting one overall number, or using the total population as the denominator for every cell, is the error.
| AG2-strat table · aggregation/stratification | not run |
| B4 | Report the annual incidence of type 2 diabetes over the last five calendar years. Report the yearly trend.Trap and rubricCorrect denominator per period: each year's incidence = new cases that year / population AT RISK that year (enrolled, not previously diabetic). The trap is using a single fixed denominator (e.g. all-time population) across years, or counting prevalent cases each year. Correct handling recomputes the at-risk denominator and excludes prior-year prevalents per calendar year. Grading: Must compute, per calendar year, new T2D cases over the at-risk enrolled population that year (excluding previously diagnosed patients). A fixed denominator across years, or prevalent counting, is the error.
| AG3-trend table · aggregation/stratification | not run |
| B5 | Is the rate of hip fracture higher in women than in men in this database? Report the comparison.Trap and rubricAge-structure confounding of a descriptive comparison: hip-fracture rate rises sharply with age and the female and male subpopulations differ in age distribution. A crude male-vs-female rate comparison conflates sex with age composition. Correct comparison age-standardizes (or stratifies by age band) before concluding. Grading: Must age-standardize the sex comparison (or compare within age strata) rather than compare crude female vs male rates. Crude comparison ignoring age structure commits this error.
| AG4-ratio table · aggregation/stratification | not run |
| B6 | Report the number of first stroke diagnoses by patient age group, where age is the patient's age at the time of the stroke.Trap and rubricAge-at-event vs age-now, and year-only birth dates: age must be computed AT the event date, not the patient's current age or age at enrollment; and when only birth YEAR is available, naive (event_year - birth_year) is off by up to a year (birthday not yet reached). A one-shot query buckets by a single stored age or current age, misclassifying patients near band boundaries. Correct handling computes age at the event date, handling year-only DOB conservatively. Grading: Must compute each patient's age AT the stroke date (not current/enrollment age) and bin into age groups, handling year-only birth dates without systematic off-by-one. Using a static/current age is the error.
| AGEEV1-ageatevent table · age-at-event | not run |
| B7 | What were the ten most common diagnoses in 2022? Rank them.Trap and rubricDistinct-patient vs record-count ranking: a patient with a chronic condition generates many condition_occurrence rows, so ranking by ROW count over-weights high-utilization conditions. 'Most common' should rank by DISTINCT patient count (each patient once per condition). A one-shot query does COUNT(*) GROUP BY condition (row grain), producing a biased ranking. Grading: Must rank conditions by COUNT(DISTINCT person_id) in 2022, each patient counted once per condition, not by condition_occurrence row count. Ranking by COUNT(*) over rows is the error.
| AGG-distinctrank table · distinct-count-ranking | not run |
| B8 | How many patients had a recent diagnosis of diabetes?Trap and rubricUndefined temporal qualifier 'recent': 'recent' has no fixed meaning — recent relative to what anchor (today, data cut-off, an index date?) and within what interval (90 days, 12 months?). A robust system FLAGS 'recent' as undefined and states or requests the interval/anchor before computing; a one-shot query silently invents a window (or ignores 'recent' and counts all diabetes), yielding an unstated, arbitrary definition. Grading rewards surfacing the ambiguity and stating the assumed definition. Grading: Must recognize 'recent' as an undefined interval, and either request the threshold/anchor or explicitly state the assumed window before computing — not silently pick an unstated window or ignore the qualifier. Silent invention is the error.
| AMB1-recent count · ambiguity-interval | not run |
| B9 | Identify unusually frequent emergency-department users.Trap and rubricUndefined threshold 'unusually frequent': 'unusually frequent' needs a defined cut (>=4 visits/year? a top percentile? over what population and period?). A robust system flags the threshold as undefined and states or requests it; a one-shot query silently picks a cut-off (or a percentile) without justification, and the population/period for 'unusual' is unstated. Grading rewards surfacing the ambiguity and declaring the operational threshold and denominator. Grading: Must recognize 'unusually frequent' as an undefined threshold, and state/request the cut-off, period, and comparison population rather than silently choosing one. Silently inventing a threshold is the error.
| AMB2-frequent count · ambiguity-threshold | not run |
| B10 | Identify patients with worsening kidney disease.Trap and rubricUndefined progression algorithm 'worsening': 'worsening' could mean an eGFR decline of a stated magnitude, a sustained/confirmed drop, a CKD-stage increase, or dialysis initiation — each a different algorithm with different confirmation rules. A robust system flags 'worsening' as undefined and states or requests the operational definition; a one-shot query silently applies one rule (often a single-value threshold) without stating it. Grading rewards surfacing the ambiguity and specifying the progression algorithm. Grading: Must recognize 'worsening' as an undefined progression construct and state/request the algorithm (eGFR decline magnitude, confirmation, stage change, dialysis) before computing, not silently apply one unstated rule. Silent single-rule choice is the error.
| AMB3-worsening count · ambiguity-algorithm | not run |
| B11 | How many patients with Parkinson's disease had a fall in 2022? Report the count.Trap and rubricAscertainment across the right code fields: falls are frequently captured only in EXTERNAL-CAUSE / injury code fields, not the primary diagnosis field. A one-shot query searches only the primary diagnosis position and undercounts falls. Correct handling searches all relevant fields (primary + secondary + external-cause positions) where a fall would be recorded, acknowledging under-ascertainment. This tests whether the system reasons about WHERE the signal lives, not just which code. Grading: Must search fall events across the appropriate fields (including external-cause/injury code positions, secondary diagnoses), not only the primary diagnosis field, acknowledging fall under-ascertainment. Searching primary diagnosis only is the error.
| ASC1-externalcause count · ascertainment-completeness | not run |
| B12 | For each patient with chronic kidney disease, report their most recent eGFR value recorded on or before January 1, 2023.Trap and rubricValue-at-max-date / greatest-n-per-group: the answer needs the eGFR VALUE associated with each patient's LATEST qualifying result date — not MAX(value), not AVG, not the latest date alone. A one-shot query commonly writes MAX(egfr) (returns the highest value, wrong row) or joins on MAX(date) without tie handling. Correct handling picks the value at the row with the maximum date per patient (ROW_NUMBER/argmax), breaking same-date ties deterministically. Grading: Must return, per patient, the eGFR value from the record with the maximum date <= 2023-01-01 (argmax on date), not MAX(value) or an average. Using MAX(value) or aggregating across dates is the error.
| ASOF1-valueatmax table · greatest-n-per-group | not run |
| B13 | Among patients newly diagnosed with rheumatoid arthritis, estimate the incidence of new-onset interstitial lung disease in the first year after the RA diagnosis. Report the count of incident patients.Trap and rubricPrevalence-vs-incidence washout, condition-anchored: patients who already had interstitial lung disease (ILD) BEFORE their RA diagnosis are prevalent, not incident, cases and must be excluded via a baseline washout/look-back. Counting every ILD diagnosis in the year after RA (including pre-existing disease surfacing in records) overcounts incidence. The washout requirement is HIDDEN. Grading: System should apply a baseline washout: exclude patients with any ILD diagnosis before (or at) the RA index date, and count only ILD first occurring in the post-index year among those with adequate prior observation. No washout (counting prevalent ILD as incident) commits this error.
| C1-washout count | not run |
| B14 | Report the monthly count of new asthma diagnoses for each month from 2019 through 2023. Report the full monthly series.Trap and rubricCalendar-series completeness (gap-filling): months with ZERO new diagnoses must still appear as 0, not be silently dropped. A one-shot GROUP BY month only emits months that have data, so zero-count months vanish and any trend/rolling calculation is silently wrong. Correct handling generates the complete month spine (2019-01.. 2023-12) and left-joins counts, filling absent months with 0. Grading: Must produce every month in the 2019-2023 range including zero-count months (generate a calendar spine and left-join). A bare GROUP BY month that omits empty months is the error.
| CAL1-zerofill table · calendar-completeness | not run |
| B15 | What is the incidence of influenza diagnoses per 1,000 person-weeks during the 2022-2023 flu season, defined as October 1, 2022 through March 31, 2023?Trap and rubricSeason-spanning person-time in weeks: the observation window crosses a calendar-year boundary, and person-time must be accrued in WEEKS across that boundary (not reset at Jan 1), with each patient's at-risk time clipped to the intersection of their enrollment and the Oct 1-Mar 31 window and censored at first flu. A one-shot query buckets by calendar year, or computes person-time in a way that resets at the year boundary or ignores partial-window enrollment. Correct handling accrues weeks continuously across the boundary. Grading: Must accrue at-risk person-WEEKS across the year-boundary-spanning Oct 1 2022-Mar 31 2023 window (clip to enrollment, censor at first flu), then events per 1,000 person-weeks. Bucketing by calendar year, or resetting person-time at the boundary, is the error.
| CALW1-personweeks rate · person-time-seasonal | not run |
| B16 | Report the number of new (incident) users of semaglutide per calendar quarter of 2022, where a new user has no semaglutide fill in the 365 days before their first fill.Trap and rubricCalendar-quarter grouping with day-based washout: the QUARTER is defined by fill date, but the 365-day washout is defined in DAYS, not calendar months. A one-shot query approximates the washout as '12 calendar months' or '4 quarters' and mis-classifies fills near quarter/month boundaries (off-by-one), or applies the washout only within the same calendar year (truncating the lookback at Jan 1). Correct handling uses an exact 365-day lookback per patient independent of the quarter bucketing. Grading: Must apply an exact 365-day (day-based) washout ending at each patient's first semaglutide drug_exposure, independent of the calendar-quarter bucket, and count only incident users per quarter. Approximating the washout in calendar months, or truncating the lookback at the year boundary, is the error.
| CALW2-washoutbound table · washout-boundary | not run |
| B17 | Assign each patient a monthly CKD stage by carrying the most recent recorded stage forward for at most 180 days; months more than 180 days after the last record are 'unknown' until a new record appears. Report the monthly stage distribution.Trap and rubricBounded state carry-forward: a recorded stage remains valid for a limited window (180 days); beyond that, status is UNKNOWN, not silently persisted forever nor dropped. A one-shot query either carries the last value forward indefinitely (overstating known status in stale months) or only reports months with an actual record (dropping carry-forward entirely). Correct handling carries forward with a 180-day expiry and emits 'unknown' months. Grading: Must carry the last recorded stage forward up to 180 days and mark later months 'unknown' until a new record, per patient. Indefinite carry-forward, or reporting only recorded months, is the error.
| CARRY1-limitedcarry table · limited-carry-forward | not run |
| B18 | What is the 1-year all-cause mortality after a patient's first heart-failure hospitalization?Trap and rubricMortality ascertainment vs administrative censoring: a patient with NO death record who DISENROLLED before day 365 is not known to be alive — their outcome is unascertained (censored), not 'survived'. A one-shot query treats every patient without a death record as a survivor, inflating the denominator with unobservable patients and biasing mortality downward. Correct handling restricts the at-risk denominator to patients observable (enrolled or with death ascertainment) through 365 days, or censors the unobserved. Grading: Must distinguish 'no death record AND observable through day 365' (survivor) from 'disenrolled before day 365 with no death record' (censored/unascertained), not count all non-death patients as survivors. Treating disenrolled-without-death as alive is the error.
| CENS1-mortascertain proportion · mortality-vs-censoring | not run |
| B19 | What is the incidence of first stroke per 1,000 person-years over 2021-2022, accounting for death as a terminating event?Trap and rubricCompeting terminating event + post-death enrollment trap: person-time must STOP at death (a patient cannot have a stroke after dying), and a patient who died in 2021 must not contribute person-time or appear in the 2022 at-risk denominator even if a stale enrollment row spans 2022. A one-shot query accrues person-time to disenrollment/window-end ignoring death, and may count dead patients as at-risk in later years. Correct handling censors person-time at death and removes decedents from subsequent denominators. Grading: Must censor at-risk person-time at death (stop accrual, exclude decedents from later-year denominators) as well as at first stroke, disenrollment, and window end. Accruing person-time past death, or keeping decedents in the 2022 denominator via a stale enrollment row, is the error.
| CENS2-competingrisk rate · competing-risk | not run |
| B20 | What is the prevalence of type 2 diabetes in 2022?Trap and rubricChronic-condition persistence vs in-year coding: diabetes is lifelong, so a patient with a condition_occurrence in 2019 but none re-coded in 2022 is still prevalent in 2022 — coding is intermittent. Requiring a diagnosis WITHIN 2022 undercounts chronic prevalence. Correct handling counts a patient as prevalent in 2022 if EVER diagnosed AND in observation during 2022 (observation_period overlaps 2022). Grading: Must count chronic (lifelong) diabetes as prevalent in 2022 based on ever-diagnosed (carry-forward) among patients in observation during 2022 (observation_period), not on the presence of a 2022-dated code. Requiring an in-2022 diagnosis undercounts chronic prevalence — the error.
| CHRON1-persistence proportion · chronic-persistence | not run |
| B21 | How many distinct healthcare-contact days did each patient have in 2022, treating multiple encounters on the same calendar date as a single contact day? Report the median across patients.Trap and rubricContact days vs encounter count: 'contact days' collapses all same-date encounters (across visit types, providers, and rows) into ONE day, so the metric is COUNT(DISTINCT visit_start_date) per patient — not the raw visit_occurrence row count. A one-shot query counts visit rows, over-counting patients with multiple same-day encounters. Grading: Must compute contact days as COUNT(DISTINCT visit_start_date) per patient, collapsing same-date visit_occurrence rows. Counting raw visit_occurrence rows is the error.
| CONT1-contactdays summary-statistic · distinct-contact-days | not run |
| B22 | How many patients have probable non-alcoholic fatty liver disease (NAFLD)? Report the count.Trap and rubricCompound phenotype with between-event exclusions: probable NAFLD = >=2 abnormal ALT results >=6 months apart AND BMI >=30, while EXCLUDING patients with viral hepatitis, alcohol use disorder, or hepatotoxic drug exposure occurring BETWEEN the two ALT measurements. A one-shot query tends to take a single ALT, ignore the 6-month spacing, or apply exclusions over all-time rather than the between-ALT interval. Correct handling composes the repeated-lab + BMI inclusion with interval-scoped exclusions. Grading: Must require >=2 abnormal ALTs >=6 months apart plus BMI>=30, and exclude patients with viral hepatitis / alcohol use disorder / hepatotoxic exposure BETWEEN the two ALTs. Single ALT, no spacing, or all-time (not interval-scoped) exclusions is the error.
| CP1-compound count · compound-phenotype-exclusions | not run |
| B23 | How many colonoscopies did each patient undergo in 2022? Report the median across patients.Trap and rubricSame-event de-duplication: one colonoscopy can generate multiple procedure_occurrence rows across encounters/sources for the same clinical event on the same day. Counting procedure rows overcounts procedures per patient. Correct handling collapses to distinct clinical events (dedup by person_id + procedure_date) before counting. A one-shot query does COUNT of procedure_occurrence rows. Grading: Must collapse multiple procedure_occurrence rows for the same colonoscopy on the same day into ONE event (dedup by person_id + procedure_date) before counting. Counting raw procedure_occurrence rows is the error.
| DEDUP1-claimline summary-statistic · claim-dedup | not run |
| B24 | Count patients with a recorded body-mass index (BMI) over 40 (severe obesity). Report the count.Trap and rubricData-quality screening before counting: BMI values in EHR data contain data-entry errors and implausible values (e.g. BMI 0, 3, or 900 from unit/typo errors). A blind count of BMI>40 also silently includes impossible values like 900 as 'severe obesity'. The correct behavior is to screen for a biologically plausible BMI range (e.g. ~10-100) before applying the >40 threshold. A one-shot query thresholds the raw field and mis-counts. Grading: System must restrict to biologically plausible BMI values (e.g. exclude <10 and >100 as data errors) before counting BMI>40. Thresholding the raw BMI field without a plausibility screen commits the error.
| DQ2-quality count · data-quality/feasibility | not run |
| B25 | How many distinct patients have a diagnosis of asthma? Report the count.Trap and rubricDuplicate/grain awareness: the DIAGNOSIS table has one row per diagnosis event, so a patient with asthma appears on many rows. A naive COUNT over the filtered table counts diagnosis ROWS (or visits), not distinct patients, massively overcounting. Correct handling counts DISTINCT patient IDs. The word 'patients' signals the required grain; the trap is answering at row grain. Grading: Must count DISTINCT patient identifiers, not diagnosis rows/events. A COUNT(*) over the filtered diagnosis table (row grain) overcounts and is the error.
| DQ3-dedup count · data-quality/feasibility | not run |
| B26 | What proportion of patients are current smokers? Report the proportion.Trap and rubricAbsence-as-negative + observability: smoking status is recorded in observation/social-history and is missing for many patients. Treating everyone without a 'current smoker' record as a non-smoker (denominator = all patients) conflates true non-smokers with unrecorded status, biasing the proportion. Correct handling restricts to patients with ANY smoking-status observation (observable), or explicitly flags the missingness, rather than assuming missing = non-smoker. Grading: Must base the proportion on patients with a recorded smoking status (observable denominator), or explicitly handle missingness — not treat absence of a smoking record as a confirmed non-smoker over the whole population.
| DQ4-active proportion · data-quality/feasibility | not run |
| B27 | How many patients have clinical activity recorded after their date of death? Report the count.Trap and rubricDeath-date temporal integrity (data-quality): a valid record cannot postdate death; encounters/diagnoses/labs/fills dated after the death date are data errors. Answering requires joining death date to all activity tables and finding any post-death record. A one-shot query rarely thinks to cross-check death against every activity source; the correct behavior is to detect the temporal impossibility. This also implicitly tests that downstream cohort logic would need to CENSOR at death. Grading: Must join each patient's date of death to activity records (encounter/diagnosis/procedure/lab/rx) and count patients with any record dated strictly after death. Ignoring death-date consistency is the error.
| DQ5-deathafter count · data-integrity | not run |
| B28 | How many laboratory results have a numeric value outside the reference range but an abnormal flag indicating normal (or vice versa)? Report the count.Trap and rubricCross-field inconsistency (value vs flag): the numeric value + reference range and the abnormal-flag column can disagree due to data-entry errors. Detecting this requires comparing the computed in/out-of-range status against the recorded flag. A one-shot query trusts one field (usually the flag) and never cross-validates. Correct handling compares value-vs-range to the flag and counts mismatches — and, importantly, this is WHY a downstream 'abnormal lab' cohort should derive abnormality from the value, not the flag. Grading: Must compare each result's numeric value against its reference range and count records where that computed status contradicts the recorded abnormal flag. Trusting a single field without cross-validation is the error.
| DQ6-valueflag count · data-integrity | not run |
| B29 | What is the incidence of hip fracture in this database? Report the incidence.Trap and rubricPrevalence-incidence + person-time: 'incidence' requires NEW fractures over person-time at risk, not a headcount of anyone with a fracture code (which conflates old/prevalent fractures and ignores unequal follow-up). A one-shot query reports a simple proportion of patients with the code. Correct handling counts incident (first) fractures and divides by person-time (or at least a properly defined at-risk denominator over a period). Grading: Must count incident (first-occurrence) hip fractures and use a person-time or period-at-risk denominator, excluding prevalent fractures. Reporting patients-with-code / total as 'incidence' commits this error.
| EM1-period rate · epi-methodology | not run |
| B30 | What is the prevalence of prostate cancer among adults? Report the prevalence.Trap and rubricSex-eligibility denominator: prostate cancer occurs (essentially) only in men, so the at-risk denominator is the male population, not all adults. A one-shot query divides by all adults, understating prevalence roughly two-fold. Correct handling restricts the denominator to men. Grading: Must restrict the denominator to the eligible (male) population. Dividing prostate-cancer cases by all adults commits this error.
| EM2-ageband proportion · epi-methodology | not run |
| B31 | What percentage of patients with diabetes have had an HbA1c test in the past year? Report the percentage.Trap and rubricObservability of the denominator: the 'in the past year' quality metric is only meaningful for patients actually enrolled/observable during that year. Including diabetics who are not enrolled in the measurement year (no chance to have a recorded test) dilutes the percentage. Correct handling restricts the denominator to diabetics observable in the measurement year. Grading: Must restrict the denominator to diabetic patients enrolled/observable during the measurement year, then compute the fraction with an HbA1c that year. Using all-ever diabetics as the denominator understates the rate.
| EM3-recentonly proportion · epi-methodology | not run |
| B32 | How many hospitalizations for sepsis occurred in 2022, and how many distinct patients did they involve? Report both.Trap and rubricEncounter-episode collapsing: inter-facility TRANSFERS produce multiple adjacent Inpatient Visits in visit_occurrence for ONE clinical hospitalization. Counting raw inpatient visit rows overcounts hospitalizations; adjacent stays (discharge and next admission within ~1 day, by visit_start_date/visit_end_date) must be collapsed into one episode. A one-shot query counts visit rows. Grading: Must collapse adjacent/overlapping Inpatient Visits (transfers, gap <= ~1 day using visit_start_date/visit_end_date) into single sepsis hospitalization episodes before counting, and separately report distinct patients. Counting raw inpatient visit rows as hospitalizations is the error.
| EP1-episode table · episode-collapsing | not run |
| B33 | Construct chronic-condition episodes and, for each, mark whether its start or end is incomplete because the episode begins before the patient's observable period starts or extends beyond when observability ceases. Report episode durations with an incomplete-boundary flag.Trap and rubricIncomplete episode boundaries / truncation: episodes clipped by the observation window are LEFT- or RIGHT-truncated — their true start/end is unknown, so their duration is a minimum, not exact. A one-shot query computes duration from observed endpoints and treats every episode as complete, biasing durations short and mis-stating incidence at window edges. Correct handling flags episodes touching the observability boundary as incomplete and treats their duration as censored. Grading: Must flag episodes whose start precedes observable-period start or whose end reaches observability cessation as boundary-incomplete (duration = minimum/censored), not treat all episodes as complete. Reporting truncated durations as exact is the error.
| EPINC1-incompleteboundary table · episode-incomplete-boundary | not run |
| B34 | Construct treatment episodes by merging overlapping supply intervals, including transitive chains where interval A overlaps B and B overlaps C so that A, B, and C form one episode even though A and C do not directly overlap. Report the median number of episodes per patient.Trap and rubricTransitive interval merging: episode construction must merge via CONNECTED COMPONENTS of overlap — if A-B and B-C overlap, all three collapse into one episode though A and C are disjoint. A one-shot query does pairwise overlap checks or self-joins that miss transitive chains, over-counting episodes. Correct handling is a gap-and-island / running-max-end sweep that closes an episode only when a true gap appears. Grading: Must merge intervals transitively (connected components / running-max-end sweep) so overlap chains form a single episode, closing only at a real gap. Pairwise-only overlap logic that splits transitive chains is the error.
| EPTR1-transitive summary-statistic · episode-transitive-merge | not run |
| B35 | Construct treatment episodes for apixaban allowing up to a 30-day gap between fills, and report the median duration of each patient's FIRST continuous treatment episode. Report the median in days.Trap and rubricDrug-era / gap-and-island construction: a continuous treatment episode chains consecutive apixaban dispensings whose gaps are <=30 days (using each record's drug_exposure_start_date..drug_exposure_end_date), and breaks when a gap exceeds 30 days. The FIRST episode's duration = its start to the end of its last contiguous exposure. A one-shot query takes last-minus-first exposure date (ignoring gaps) or counts exposures. Correct handling stitches drug_exposure rows into gap-bounded episodes per patient and measures the first. Grading: Must build gap-bounded treatment episodes (consecutive apixaban drug_exposure rows with <=30-day gaps, using drug_exposure_start_date/end_date), identify each patient's first episode, compute its duration, then take the median across patients. Using last-minus-first exposure (ignoring gaps) or exposure counts is the error.
| ER1-drugera summary-statistic · drug-era-gap | not run |
| B36 | How many patients met all inclusion criteria for an anticoagulation cohort but were excluded by exactly one exclusion criterion? Report the count and which single exclusion it was.Trap and rubricExclusion accounting at the patient level: the question needs patients who satisfy ALL inclusions and fail EXACTLY ONE of several exclusions — requiring a per-patient count of how many exclusions each triggers, then keeping those with count==1. A one-shot query applies exclusions as a combined NOT filter (removing anyone failing any exclusion) and cannot report how many patients failed exactly one, nor which. Correct handling evaluates each exclusion independently per patient and counts the triggers. Grading: Must, among inclusion-meeting patients, count exclusions triggered per patient and keep those triggering exactly one (reporting which). Applying exclusions as a single combined filter (no per-exclusion count) is the error.
| EXCL1-oneexclusion table · exclusion-accounting | not run |
| B37 | What is the prevalence of atrial fibrillation among adult patients? Report the prevalence.Trap and rubricActive-vs-historical status discovery: condition_occurrence carries condition_status_concept_id, whose values distinguish active/primary diagnoses from 'History of' (resolved/past) records. Prevalence of a current condition must exclude 'History of'. A one-shot query written blind counts every condition_occurrence row, inflating prevalence. The status field and its values are discoverable only by inspecting the table/vocabulary — the prompt gives no hint. Grading: System should discover condition_status_concept_id and exclude 'History of' records (count active/primary diagnoses). Requires inspecting the table before writing the count. A blind query over all atrial-fibrillation condition_occurrence rows commits this error by conflating active and historical disease.
| F1-dxstatus proportion · flexibility/runtime-discovery | not run |
| B38 | What is the mean age of patients with a heart failure diagnosis? Report the mean age.Trap and rubricImplausible-value discovery: patient birth-year/age fields contain sentinel and implausible values (e.g. birth year 1900, age 0, or ages >120 from data-entry errors). A blind AVG over the raw age column is skewed by these outliers. The correct behavior is to inspect the age distribution, recognize the implausible values, and filter them (e.g. 18-100) before averaging. A one-shot AVG reports the skewed number. The bad values are visible only on inspection. Grading: System should inspect the age/birth-year distribution, detect implausible/sentinel values, and exclude them (plausible adult range) before computing the mean. A raw AVG over unfiltered ages commits this error.
| F5-sentinel summary-statistic · flexibility/runtime-discovery | not run |
| B39 | Count patients who were hospitalized (had an inpatient stay) in the last year. Report the count.Trap and rubricColumn/value disambiguation: 'inpatient' is encoded in a specific VISIT field/value (e.g. visit_type or a care-setting code) whose exact coding is not obvious. A blind query guesses a value (e.g. visit_type='IP') that may not match the actual encoding, silently returning wrong counts. The correct behavior is to inspect the visit table's sample values to find how inpatient stays are actually flagged, then filter on the real value. Grading: System should inspect the VISIT table to discover how inpatient/hospitalization is actually encoded (sample the relevant column's values), then filter on the correct value. Guessing a code without verifying against the data risks matching nothing or the wrong rows.
| F6-inpatient count · flexibility/runtime-discovery | not run |
| B40 | For patients with diabetes, report the total number of outpatient visits, broken down by whether the patient also has hypertension.Trap and rubricJoin fan-out inflation: joining a patient's visits (one-to-many) to their diagnoses (also one-to-many) produces a Cartesian fan-out, so COUNT/SUM over the joined rows multiplies visits by the number of matching diagnosis rows. A one-shot query joins the tables then counts, inflating visit totals. Correct handling aggregates visits at the correct grain FIRST (distinct visits per patient) and joins the comorbidity flag separately, avoiding many-to-many multiplication. Grading: Must avoid many-to-many fan-out: count DISTINCT visits (aggregate to patient/visit grain before or independent of the diagnosis join), then attach the hypertension flag. Counting rows of a diagnosis-to-visit join is the error (visits multiplied by matching diagnosis rows).
| FAN1-joinfanout table · join-fan-out | not run |
| B41 | Among continuously enrolled patients, what is the median longest gap (in days) between consecutive healthcare encounters? Report the median.Trap and rubricMax inter-event gap per patient: requires ordering each patient's encounters, computing gaps between consecutive encounters, taking the MAXIMUM gap per patient, then the median of those maxima. A one-shot query tends to compute average gaps, or last-minus-first, not the per-patient maximum consecutive gap. Correct handling uses ordered lead/lag differences per patient then aggregates the per-patient maxima. Grading: Must order encounters per patient, compute consecutive-encounter gaps, take each patient's MAX gap, then median across patients. Using average gaps or total span is the error.
| GAP1-maxgap summary-statistic · max-inter-event-gap | not run |
| B42 | For each month of 2021, report how many patients were in an active diabetes cohort as of that month, using only information available on or before the end of that month.Trap and rubricNo-lookahead / point-in-time correctness: each month's membership must be reconstructed using ONLY records dated on or before that month-end — a later diagnosis or the patient's final/current disease status must not leak backward into earlier months. A one-shot query applies the patient's overall (ever/current) status to every month, contaminating historical counts with future information. Correct handling evaluates membership as-of each month using only then-available records. Grading: Must determine each month's cohort membership using only records available on or before that month-end (no future/current-status leakage into past months). Applying an ever/current flag uniformly across months is the error.
| HIST1-nolookahead table · point-in-time-correctness | not run |
| B43 | For each patient, select one index diabetes diagnosis using the encounter-type priority inpatient > emergency > outpatient > unspecified, applied among the earliest qualifying records. Report the count of index records by encounter type.Trap and rubricIndex selection by encounter-type hierarchy: when several qualifying diagnoses share the earliest date, the index must be chosen by a stated encounter-type PRIORITY, not arbitrarily — and exactly one index per patient. A one-shot query takes MIN(date) and, on ties, returns multiple rows (double-count) or an arbitrary row, ignoring the clinical hierarchy. Correct handling ranks earliest-date records by the priority and picks one deterministically. Grading: Must pick exactly one index per patient among earliest-date records using the encounter-type priority (inpatient>ED>outpatient>unspecified), breaking residual ties deterministically. A MIN(date) selection that returns multiple/arbitrary rows on ties is the error.
| IDX1-hierarchy table · index-encounter-hierarchy | not run |
| B44 | For a two-record phenotype requiring two qualifying diagnoses 30 to 180 days apart, determine each qualifying patient's index date as the date of the earliest first-record that has a valid confirming second record, resolving ties by encounter type then record identifier. Report the count.Trap and rubricEarliest valid confirming pair: the index is the earliest FIRST record that actually has a partner 30-180 days later — not simply the earliest diagnosis (which may lack a valid confirmation) nor the second record. Requires, per patient, searching for the earliest first-record with a qualifying second in the window, with deterministic tie-breaking. A one-shot query takes the first two diagnoses or MIN(date) without validating the 30-180-day pairing, mis-dating the index. Grading: Must find the earliest first-record possessing a confirming second record 30-180 days later (deterministic tie-break), using that first-record's date as index. Using the earliest diagnosis regardless of confirmation, or the second record, is the error.
| IDX2-confirmingpair count · index-confirming-pair | not run |
| B45 | How many patients' first recorded stroke event changes depending on whether the record-line date, the encounter start date, or the admission date is used to order events? Report the count of patients whose index date differs across these fields.Trap and rubricDate-field sensitivity for index selection: a single clinical event can carry several dates that disagree — condition_start_date on condition_occurrence versus visit_start_date (admission) on the associated visit_occurrence — so the 'first' event, and thus the index date, can shift with the chosen field. A one-shot query picks whichever date is handy, unaware the choice moves the cohort. Grading: Must compute each patient's first stroke under condition_start_date versus visit_start_date (admission) definitions and count patients whose index differs across them, demonstrating awareness that the date field is a modeling choice. Silently using one date without acknowledging the sensitivity is the error.
| IDX3-datefield count · date-field-sensitivity | not run |
| B46 | How many patients with diabetes are on metformin? Report the count.Trap and rubricCurrent-use vs ever-use ambiguity: 'on metformin' is ambiguous between CURRENT use (active drug_exposure interval covering a reference date) and EVER use (any prior exposure). The two differ greatly. A one-shot query silently picks 'any exposure ever' without stating the choice or building an active-coverage timeline from drug_exposure_start_date/end_date. Grading: Must adopt and STATE a use definition; for current use, determine active coverage (drug_exposure interval covering the reference date), not merely any prior exposure. Silently equating 'on metformin' with 'ever had a metformin exposure' (undeclared) is the error.
| INT1-currentvsever count · current-vs-ever | not run |
| B47 | What is the average number of emergency-department visits per patient in 2022?Trap and rubricZero-inclusive denominator: 'per patient' should average over ALL enrolled patients, including the many with ZERO ED visits, not only over patients who had at least one visit. A one-shot query computes total visits / patients-with-a-visit (implicitly dropping zeros), producing a figure that can be an order of magnitude too high. Correct handling divides total ED visits by the full enrolled denominator (zeros included), and states it. Grading: Must divide total ED visits by ALL enrolled patients (including those with zero visits) for a per-enrollee average, not by only patients with >=1 visit. Restricting the denominator to visit-having patients is the error.
| INT2-zerodenominator summary-statistic · zero-inclusive-denominator | not run |
| B48 | What share of amoxicillin prescriptions were for a viral upper-respiratory infection (i.e. potentially inappropriate)? Report the share.Trap and rubricExposure-to-indication episode linkage: each amoxicillin drug_exposure must be LINKED to a URI condition_occurrence in the same clinical episode (within +/-3 days) AND have no competing bacterial indication in that window — a per-EXPOSURE linkage, not a patient-level intersection. A one-shot query intersects 'patients with amoxicillin' and 'patients with a URI' anytime. Grading: Must link each amoxicillin drug_exposure to a URI condition_occurrence within +/-3 days (same episode) with no competing bacterial indication in that window, then compute the share of exposures so linked. Intersecting patient lists (amoxicillin-ever AND URI-ever) is the error.
| INT3-indicationlink proportion · indication-linkage | not run |
| B49 | What is the most-prescribed medication among patients aged 65 and older in 2022?Trap and rubricRanking-metric ambiguity + brand/generic normalization: 'most-prescribed' has multiple defensible metrics — by fill count, by distinct patients, by total days-supply, or by unique molecule after normalizing brand/generic names — which can yield DIFFERENT top drugs. A one-shot query silently ranks by raw fill rows without normalizing brand and generic to one ingredient (so the same molecule splits across names and loses). Correct handling states the metric and normalizes brand/generic to the ingredient level before ranking. Grading: Must state which 'most-prescribed' metric is used (fills / distinct patients / days-supply / molecules) and normalize brand+generic to the ingredient level before ranking. Ranking raw fill rows without brand/generic normalization or metric declaration is the error.
| INT4-rankmetric table · ranking-metric-ambiguity | not run |
| B50 | For each patient, what is the total length of continuous observable time, merging enrollment periods separated by gaps of no more than 30 days? Report the median across patients.Trap and rubricInterval-merge primitive: a patient's enrollment is stored as multiple spans; continuous observable time requires MERGING spans that overlap or are <=30 days apart into consolidated intervals, then summing. A one-shot query sums raw span lengths (double-counting overlaps) or uses last-minus-first (counting gaps as observed). Correct handling is interval coalescing (gap-and-island on date ranges). Grading: Must coalesce enrollment spans that overlap or are <=30 days apart into merged intervals, then sum their lengths per patient (median across patients). Summing raw spans (overlap double-count) or last-minus-first (includes true gaps) is the error.
| IV1-mergeobs summary-statistic · interval-merge | not run |
| B51 | Among metformin initiators, estimate the incidence of new-onset diabetic retinopathy in the first year after initiation. Report the count of incident patients.Trap and rubricPrevalent-vs-incident washout: prevalent distractors have retinopathy BEFORE metformin AND a follow-up visit at days 30-45. A baseline lookback washout (exclude patients with prior retinopathy) correctly identifies them. CLEAN also tests over-washing: a fixed induction period (>90d) wrongly drops true early cases at days 15-29 that have no prior history. Grading: System must apply a baseline lookback washout (exclude patients with retinopathy before metformin start). This gives gold on both scenarios. A fixed induction period (>90d) gives gold on FIRES but over-washes on CLEAN (drops true early cases at days 15-29 with no prior history).
| L3-056 count | not run |
| B52 | Among oral-contraceptive users, count the patients who had a genuine, clinically confirmed pulmonary embolism during follow-up.Trap and rubricRule-out / low-specificity outcome coding: emergency-department pulmonary-embolism (PE) diagnosis codes are frequently written to justify a CT scan even when PE is ruled out. Counting every PE code massively overcounts. A genuine PE needs multi-source triangulation: the diagnosis PLUS a confirmatory imaging procedure (CT pulmonary angiography) PLUS therapeutic anticoagulation. This is goal-only: the confirmation requirement is NOT stated in the prompt. Grading: System must confirm PE by triangulating three entity types — PE diagnosis + confirmatory CTPA imaging procedure + therapeutic anticoagulant — rather than counting the PE diagnosis code alone. Counting any PE code (or a single-source proxy) commits this error (overcount from rule-out cases).
| L3-057 count | not run |
| B53 | Measure the rate of acute kidney injury among patients with severe (Stage 4/5) chronic kidney disease. Make sure the CKD cohort reflects true disease severity.Trap and rubricSeverity under-coding (labs > codes): physicians under-report CKD severity using generic CKD codes, so relying on Stage 4/5 billing codes (N18.4/N18.5) severely undercounts true severe CKD. True severity must be established physiologically from raw serum creatinine via the CKD-EPI eGFR formula (eGFR < 30 = Stage 4/5). This is NOT recoverable by better medical coding (the correct codes are absent) and requires a non-SQL computation. Grading: System must build the severe-CKD cohort from raw serum creatinine + age + sex by computing eGFR (CKD-EPI) and thresholding eGFR < 30, NOT from CKD stage billing codes alone. Relying on Stage 4/5 codes commits this error (undercount). A one-shot text-to-SQL query cannot compute CKD-EPI; the eGFR computation skill must be invoked.
| L3-059 proportion | not run |
| B54 | Estimate the rate of device-related (systemic) infection following pacemaker implantation over one year. Exclude early post-surgical infections within 30 days of implantation (surgical-site, not device-related endocarditis).Trap and rubricTemporal exclusion / pathophysiologic distinctness: early infections (0-30 days) are surgical-site (sterile technique), not device-related systemic infection. Pooling both biases the device safety profile. Grading: System must exclude infections in days 0-30 and count only late (31-365) device-related infections. It must NOT over-exclude (e.g. a 60/90-day cut drops legitimate day 31-45 cases). Pooling all 0-365 commits the error.
| L3-060 proportion | not run |
| B55 | Among patients undergoing colonoscopy, count those whose colonoscopy was diagnostic rather than routine preventive screening. Treat a colonoscopy as diagnostic only if a symptom (abdominal pain or GI bleeding) is documented within 30 days before the procedure.Trap and rubricIntent disambiguation: colonoscopies are mostly preventive. Without screening codes, treating all colonoscopies as diagnostic overcounts symptomatic cases. Diagnostic intent needs a documented prior symptom. Grading: System must require a symptom (abdominal pain or GI bleeding) documented within 30 days before the colonoscopy, and respect the 30-day window (not count symptoms 45-60d prior). Counting all colonoscopies as diagnostic commits this error.
| L3-064 count | not run |
| B56 | Among patients on chronic anticoagulation, count those who are adherent (medication coverage >= 80% of the year). Ensure inpatient hospital stays are not misread as gaps in medication.Trap and rubricExposure misclassification via inpatient coverage gaps: while hospitalized (an Inpatient Visit in visit_occurrence) patients receive medications from the inpatient formulary that do not appear as outpatient drug_exposure records, so an outpatient-only coverage calculation counts inpatient stays as medication gaps and underestimates adherence. Inpatient length-of-stay days (visit_start_date..visit_end_date) must be credited as covered in the timeline. Grading: System must detect Inpatient Visits (length of stay from visit_start_date/visit_end_date) and credit those days as medication-covered before computing coverage/adherence (>= 80%). A pure outpatient drug_exposure coverage calculation commits this error by treating hospital days as gaps and understating adherence.
| L3-066 count | not run |
| B57 | Count glaucoma patients with confirmed disease progression on visual-field testing.Trap and rubricOutcome confirmation / test noise: a single worsening visual-field test is often subjective noise. Confirmed progression needs two consecutive worsening tests >=30 days apart. The confirmation requirement is HIDDEN — the system must discover it. Grading: System should discover that a single worsening test is unreliable and require two consecutive worsening visual-field tests at least 30 days apart. Counting any single worsening test commits this error.
| L3-067-goal count | not run |
| B58 | Count patients with a true Type 1 (atherothrombotic) myocardial infarction. Define Type 1 as: a primary-position MI diagnosis AND troponin > 10x the upper limit of normal; exclude patients with a concurrent sepsis diagnosis and only mildly elevated troponin (Type 2 demand ischemia).Trap and rubricOutcome misclassification / MI subtype: sepsis causes Type 2 demand ischemia (troponin leak) that is clinically distinct from Type 1 plaque rupture. Pooling any MI billing code sweeps in Type 2 demand cases, biasing true coronary rates. Grading: System must combine three signals: primary-position MI diagnosis AND troponin > 10x ULN AND exclude concurrent sepsis with only mildly elevated troponin. Counting any MI code commits this error.
| L3-069 count | not run |
| B59 | Count patients with a true (atherothrombotic) myocardial infarction.Trap and rubricOutcome misclassification / MI subtype: a generic MI code pools true Type 1 (plaque rupture) with Type 2 demand ischemia (sepsis-driven troponin leak). A high-specificity Type 1 cohort needs primary-position dx + high troponin and sepsis exclusion. The subtype distinction is HIDDEN — the system must discover it. Grading: System should discover that MI codes conflate Type 1 and Type 2, and restrict to true atherothrombotic MI (e.g. primary-position dx + troponin >10x ULN, excluding sepsis + mildly elevated troponin). Counting any MI code commits this error.
| L3-069-goal count | not run |
| B60 | Build a cohort of patients who are NEW (incident) users of atorvastatin, and report its size. A patient already being treated for the same condition with a related medication is not a new user.Trap and rubricPrevalent-user bias via class switching: patients switching from another statin or lipid-lowering drug to atorvastatin look drug-incident but are class-prevalent. A drug-specific washout (atorvastatin only) misses these switchers. The washout must span the whole antilipemic drug class (ATC/AHFS), not just atorvastatin. Goal-only: the class-wide requirement is not stated explicitly. Grading: System must apply a class-wide washout over all antilipemic/lipid-lowering drugs (not just atorvastatin) when defining incident atorvastatin users, excluding prior users of any drug in the class. A drug-specific (atorvastatin-only) washout, or no washout, commits this error by counting class-prevalent switchers as incident.
| L3-071 count | not run |
| B61 | Count patients with a genuine Clostridioides difficile infection (CDI). Require a positive stool PCR or toxin test - do not rely on the CDI diagnosis alone (codes are often applied before testing returns).Trap and rubricRule-out / low-specificity coding: hospitalized diarrhea is often coded as CDI before the stool test returns. Counting CDI codes without lab confirmation overcounts. A genuine case needs a positive stool PCR/toxin test. Grading: System must join the CDI diagnosis to a positive stool PCR/toxin result and exclude negative/absent tests. Counting any CDI code commits this error.
| L3-072 count | not run |
| B62 | Among PPI users, count the patients with a genuine Clostridioides difficile infection (CDI).Trap and rubricRule-out / low-specificity coding: CDI is often coded on hospitalized diarrhea before the stool test returns. A genuine case needs a positive stool PCR/toxin test. The confirmation requirement is HIDDEN — the system must discover that codes overcount and lab confirmation is required. Grading: System should discover that CDI codes over-capture and require a positive stool PCR/toxin test (excluding negative/absent tests). Counting any CDI code commits this error.
| L3-072-goal count | not run |
| B63 | Count patients with true chronic End-Stage Renal Disease (ESRD) among dialysis recipients. Require dialysis spanning >= 90 consecutive days - temporary dialysis for acute, recoverable kidney injury should not qualify.Trap and rubricAcute-vs-chronic duration: ICU patients with AKI receive temporary CRRT but recover. Counting any dialysis procedure code misclassifies temporary ICU dialysis as chronic ESRD. Grading: System must compute dialysis span/continuity and require >= 90 consecutive days, respecting the boundary (30-89 day spans fail). Counting any dialysis code commits this error.
| L3-075 count | not run |
| B64 | Identify acute pancreatitis hospitalizations. Require a lipase or amylase measurement > 3x the laboratory upper limit of normal - mild chronic elevations (chronic insufficiency / renal clearance) do not qualify.Trap and rubricAcute-vs-chronic lab threshold: amylase/lipase > 3x ULN defines acute pancreatitis; mild chronic elevations represent chronic insufficiency or renal clearance. Trusting any pancreatitis code overcounts mild/chronic cases. Grading: System must require a lipase or amylase value > 3x the lab ULN, joined to the hospitalization. Counting any pancreatitis code commits this error.
| L3-083 count | not run |
| B65 | Identify acute pancreatitis hospitalizations. Report the count.Trap and rubricAcute-vs-chronic lab threshold: a pancreatitis code sweeps in mild chronic elevations. True acute pancreatitis needs lipase/amylase > 3x ULN. The lab-threshold requirement is HIDDEN — the system must discover it. Grading: System should discover that a pancreatitis code over-captures and require a lipase or amylase > 3x ULN. Counting any pancreatitis code commits this error.
| L3-083-goal count | not run |
| B66 | Identify patients with chronic non-cancer opioid dependency, and report the cohort size.Trap and rubricIndication-defining exclusion (cohort restriction): chronic opioid therapy in cancer patients is palliative and pathophysiologically distinct from chronic non-cancer dependency. To build a NON-CANCER dependency cohort, cancer patients must be excluded. This is goal-only and adversarial: the word "non-cancer" is in the prompt but the exclusion step is NOT stated — the system must discover that it needs to identify and remove oncology patients. Grading: System must exclude patients with an oncology diagnosis from the chronic-opioid cohort before counting, to isolate non-cancer dependency. Counting all chronic opioid users (cancer + non-cancer) contaminates the cohort with palliative use.
| L3-085 count | not run |
| B67 | Identify hospitalizations for true diabetic ketoacidosis (DKA). Require lab-confirmed metabolic acidosis: serum bicarbonate < 18 mEq/L AND anion gap > 12 - simple hyperglycemia billed as DKA does not qualify.Trap and rubricConjunctive lab confirmation: true DKA is defined by metabolic acidosis (bicarbonate < 18 mEq/L, anion gap > 12); simple hyperglycemia billed as DKA lacks acidosis. A DKA code alone overcounts. Grading: System must require BOTH bicarbonate < 18 mEq/L AND anion gap > 12, joined to the hospitalization. A DKA code alone, or only one lab criterion, commits this error.
| L3-086 count | not run |
| B68 | Identify hospitalizations for true diabetic ketoacidosis (DKA). Report the count.Trap and rubricConjunctive lab confirmation: simple hyperglycemia is often billed as DKA. True DKA needs metabolic acidosis (bicarbonate < 18 mEq/L AND anion gap > 12). The lab requirement is HIDDEN — the system must discover it. Grading: System should discover that DKA codes over-capture and require lab-confirmed metabolic acidosis (bicarbonate < 18 AND anion gap > 12). A DKA code alone commits this error.
| L3-086-goal count | not run |
| B69 | Identify febrile neutropenia hospitalizations in oncology patients. Because febrile neutropenia is under-coded, define it by a measured absolute neutrophil count < 1,000 cells/uL within 24h of a fever hospitalization - do not rely on the febrile-neutropenia code alone.Trap and rubricSeverity under-coding (labs > codes): febrile neutropenia is under-coded; relying on the D70.1 code severely undercounts. True cases are recovered from ANC < 1,000 labs joined to a fever admission. Grading: System must recover cases from a measured ANC < 1,000 cells/uL within 24h of a fever hospitalization, not the febrile-neutropenia code alone. Code-only commits this error (undercount).
| L3-092 count | not run |
| B70 | Identify patients with true iron-deficiency anemia. Require a ferritin < 30 ng/mL - anemia of chronic disease (ferritin typically >= 100 ng/mL) must be excluded even when coded as iron deficiency.Trap and rubricLab-threshold phenotyping: iron-deficiency anemia requires ferritin < 30 ng/mL; anemia of chronic disease has ferritin >= 100. A generic anemia code sweeps in chronic-disease anemia. Grading: System must require a ferritin < 30 ng/mL measurement and exclude chronic-disease anemia. Counting any anemia code commits this error.
| L3-095 count | not run |
| B71 | Identify patients with true iron-deficiency anemia. Report the count.Trap and rubricLab-threshold phenotyping: a generic anemia code sweeps in anemia of chronic disease. True iron-deficiency anemia needs ferritin < 30 ng/mL. The ferritin requirement is HIDDEN — the system must discover it. Grading: System should discover that anemia codes over-capture and require ferritin < 30 ng/mL, excluding anemia of chronic disease (ferritin >= 100). Counting any anemia code commits this error.
| L3-095-goal count | not run |
| B72 | Count patients with elevated natriuretic peptide indicating heart failure. Values come from two assays - BNP and NT-proBNP (by LOINC). Apply BNP > 100 pg/mL OR NT-proBNP > 300 pg/mL - not one threshold for both.Trap and rubricAssay harmonization: NT-proBNP and BNP are different assays; NT-proBNP runs ~3-4x higher for the same heart-failure severity. Applying one threshold to both misclassifies. Grading: System must split by assay (LOINC) and apply the correct threshold to each: BNP > 100 pg/mL, NT-proBNP > 300 pg/mL. A single threshold on all values commits this error.
| L3-103 count | not run |
| B73 | Count patients with definite infective endocarditis. Require BOTH positive blood cultures AND an echocardiogram - a diagnosis code alone is insufficient (many are suspected/rule-out).Trap and rubricConjunctive composite confirmation: definite infective endocarditis (Duke criteria) requires positive blood cultures and an echocardiogram showing vegetation. Many endocarditis codes are suspected/rule-out. Grading: System must require BOTH a positive blood culture AND an echocardiogram procedure, joined to the endocarditis diagnosis. A code alone, or only one component, commits this error.
| L3-110 count | not run |
| B74 | Count patients with a true recurrence of C. difficile. A recurrence is a second positive test/course 14-56 days after the first - <14 days is the same episode; >56 days is a new infection.Trap and rubricRecurrence-window misspecification (temporal): a true CDI recurrence occurs 14-56 days after the first episode. Counting any second episode conflates same-episode continuation (<14d) and re-infection (>56d) with true recurrence. Grading: System must compute the inter-episode interval and count only second episodes 14-56 days after the first. Counting any second episode (ignoring both bounds) commits the temporal error.
| L3-130 count | not run |
| B75 | Build an Ankylosing Spondylitis cohort. Report the count.Trap and rubricUnder-specification / low-specificity coding: ankylosing spondylitis is frequently miscoded from generic inflammatory back pain. A high-specificity cohort must be confirmed by HLA-B27 positivity and absence of rheumatoid factor. Counting any AS diagnosis code overcounts. The confirmation requirement is NOT stated in the prompt — the system must discover it. Grading: System should confirm AS beyond the diagnosis code — e.g. require HLA-B27 positivity and/or exclude seropositive (RF+) patients who are more likely to have a different arthropathy. Counting any AS code without lab confirmation over-captures.
| L3-143 count | not run |
| B76 | Count active acromegaly patients. Report the count.Trap and rubricAge-adjusted lab normalization: active acromegaly requires an elevated IGF-1 above the AGE-SPECIFIC upper limit of normal (IGF-1 falls with age, so a single flat cutoff misclassifies across ages). Counting any acromegaly code, or applying a flat IGF-1 threshold, mis-captures. The age-adjustment requirement is HIDDEN — the system must discover that IGF-1 reference ranges are age-dependent. Grading: System should confirm acromegaly with IGF-1 above the age-adjusted upper limit of normal (per-age reference), not a flat cutoff and not the diagnosis code alone. A flat IGF-1 threshold, or code-only, commits this error.
| L3-144 count | not run |
| B77 | Among patients with at least two blood-pressure readings per year for three consecutive years, what proportion had controlled blood pressure (<140/90) in ALL three years versus in ANY year? Report both proportions.Trap and rubricPer-patient longitudinal aggregation vs row-level filtering: 'controlled in ALL years' requires aggregating each patient's readings PER YEAR, deriving a per-year control flag, then requiring all three years true — a patient-level operation. A one-shot query typically filters ROWS where BP<140/90 (row-level), which conflates 'has some controlled reading' with 'controlled all year' and cannot express ALL-years vs ANY-years. Correct handling groups per patient-year, then reasons across years per patient. Grading: Must aggregate to a per-patient-per-year control status (e.g. all/most readings <140/90 that year), then compute (a) fraction of patients controlled in every one of the 3 years and (b) fraction controlled in at least one year — patient-level, not row-level. A row-level BP<140/90 filter is the error.
| LA1-longitudinal table · longitudinal-aggregation | not run |
| B78 | For a diabetic cohort, report the median HbA1c value closest to each patient's index date (within +/- 90 days), using exactly one value per patient. Report the median.Trap and rubricWindowed dedup with tie-break (one row per patient): each patient may have many HbA1c results near index; the metric needs the SINGLE result closest to index within +/-90 days, ties broken by most recent, then a median across patients. A one-shot query tends to take all in-window results (many per patient) and median over ROWS, over-weighting patients with more tests. Correct handling ranks per patient by |date-index| (tie-break recency), keeps the top one, then medians across patients. Grading: Must select, per patient, the single HbA1c closest to index within +/-90 days (tie-break: most recent), then take the median of those one-per-patient values. Medianing all in-window results (multiple per patient) is the error.
| LA2-indexlab summary-statistic · longitudinal-aggregation | not run |
| B79 | Among patients with at least four eGFR measurements over at least two years, how many had a sustained decline of 40% or more from baseline, confirmed by a second measurement at least 90 days after the first qualifying low value? Report the count.Trap and rubricConfirmed sustained change, not a single crossing: a >=40% decline must be CONFIRMED by a second low measurement >=90 days later, so a lone transient low value does not qualify. A one-shot query flags any single eGFR that is >=40% below baseline (counting transient dips / lab error) and ignores the confirmatory-persistence requirement. Correct handling requires two qualifying low values >=90 days apart relative to a properly defined baseline. Grading: Must require a >=40% drop from baseline CONFIRMED by a second qualifying measurement >=90 days after the first, per patient (not a single low value). Counting any single sub-threshold measurement is the error.
| LAB2-sustaineddecline count · confirmed-lab-change | not run |
| B80 | Report the median baseline creatinine, defined as the measurement closest to the index date within the window from 180 days before to 7 days after index (inclusive of day +7).Trap and rubricAsymmetric baseline window with boundary inclusivity: the baseline window is ASYMMETRIC (-180 to +7 days) and boundary-inclusive on the +7 side; 'closest to index' means the single nearest value (by absolute day distance), not the earliest, latest, or an average. A one-shot query uses a symmetric window, takes the pre-index or first value, or averages all in-window values. Correct handling selects the one measurement with minimum |days from index| within (-180, +7]. Grading: Must select, per patient, the single creatinine closest to index (minimum absolute day distance) within the asymmetric -180 to +7 day window (respecting boundary inclusivity), then median across patients. Symmetric windows, first/last value, or averaging in-window values is the error.
| LAB3-baselinewindow summary-statistic · baseline-window | not run |
| B81 | How many patients have a sequence of three consecutive laboratory results with strictly increasing numeric values? Report the count.Trap and rubricStrictly increasing consecutive triple: requires ordering each patient's results by date and finding three CONSECUTIVE measurements with strictly rising values — an ordered-sequence property, not 'has three results' nor 'max > min'. A one-shot query checks that a high and a low value exist, or counts patients with >=3 results, ignoring order and consecutiveness. Correct handling uses ordered lag comparisons over adjacent results. Grading: Must order results by date and detect three CONSECUTIVE strictly increasing values per patient (adjacent-triple comparison). Checking only that >=3 results exist, or max>min, is the error.
| LABINC1-increasingtriple count · monotonic-lab-run | not run |
| B82 | Construct periods of laboratory abnormality that begin with the first abnormal result and end with the first subsequent normal result, and report the median duration of these abnormal periods.Trap and rubricAbnormal-period construction: an abnormal period runs from an abnormal result until the NEXT normal result (which closes it); a patient may have several such periods, and a still-open period (no subsequent normal) is censored. A one-shot query measures span between first and last abnormal (ignoring interspersed normals) or treats each abnormal result independently. Correct handling walks the ordered results, opening on abnormal and closing on the first following normal. Grading: Must build abnormal periods that open at an abnormal result and close at the first subsequent normal (multiple periods per patient; open periods censored), then median their durations. Using first-to-last-abnormal span, ignoring intervening normals, is the error.
| LABPER1-abnormalperiod summary-statistic · abnormal-lab-period | not run |
| B83 | What proportion of emergency-department visits resulted in an inpatient admission on the same or the next calendar day? Report the proportion.Trap and rubricEncounter linkage (ED -> admission): each ED visit must be linked to a subsequent inpatient admission occurring same-day or next-day for that patient. A one-shot query tends to check 'patient had an ED visit AND an inpatient stay' anytime, or joins without the same/next-day constraint, overcounting. Correct handling links each ED visit to an admission within the 0-1 day window per patient. Grading: Must link each ED visit to an inpatient admission on the same or next calendar day (per patient) and compute the proportion of ED visits so linked. Any-time ED+inpatient co-occurrence is the error.
| LNK1-edadmit proportion · encounter-linkage | not run |
| B84 | Report the distribution of patients by race/ethnicity category.Trap and rubricMissing-category handling in a distribution: patients with missing/unknown race must be shown as their own category, not silently dropped — dropping them shrinks the denominator and inflates the percentages of the observed categories, misrepresenting the population. A one-shot query filters out NULL/unknown (or ignores it in the GROUP BY), so the reported percentages sum over a reduced base. Correct handling includes a 'missing/unknown' bucket and reports it against the full population denominator. Grading: Must include a missing/unknown category and compute percentages against the FULL population (missing not dropped from the denominator). Silently excluding missing values (inflating other categories' shares) is the error.
| MISS1-missingcategory table · missing-in-denominator | not run |
| B85 | How many diabetes patients completed individual monitoring components (HbA1c, lipid panel, eye exam, nephropathy screen) but never had ALL of them completed within a single 365-day window? Report the count.Trap and rubricMonitoring-bundle completeness within a window: completing the bundle requires ALL components inside ONE rolling 365-day window; a patient may have every component at some point yet never all together in 365 days. A one-shot query checks each component ever (ANDing lifetime presence), counting such patients as complete. Correct handling searches for a 365-day window containing all components and flags patients who have each individually but never concurrently. Grading: Must require all bundle components within a single rolling 365-day window to count as complete, and identify patients with each component present individually but never all within one 365-day window. ANDing lifetime component presence is the error.
| MON1-bundle count · monitoring-bundle-completeness | not run |
| B86 | Among patients who underwent total knee replacement, how many were later diagnosed with a prosthetic joint infection AND subsequently underwent a revision surgery? Report the count.Trap and rubricCompositional multi-step cohort: the answer is an INTERSECTION of three sequential sub-cohorts — (1) knee-replacement procedure, (2) later prosthetic-joint-infection diagnosis, (3) still-later revision procedure. A one-shot pipeline tends to flatten this into a single filter and loses the sequential dependency (e.g. counts anyone with all three codes regardless of order/linkage). Correct handling requires decomposing into steps and chaining them per patient. Grading: System must build the three sub-cohorts and intersect them per patient with the correct sequence (replacement -> later infection -> later revision), not just count patients who have all three codes anywhere. Flattening to a single co-occurrence filter is the error.
| MS1-comp count · multi-step/compositional | not run |
| B87 | Of patients hospitalized for heart failure, what fraction were readmitted for heart failure within 30 days of discharge? Report the fraction.Trap and rubricCompositional cohort with linked denominator: denominator = index HF hospitalizations; numerator = those with a SECOND HF hospitalization whose admission is within 30 days of the index discharge. Requires linking each index stay to its own discharge date and searching a per-patient window. A one-shot query that counts 'patients with >=2 HF admissions' ignores the 30-day linkage and the index/readmission pairing. Grading: System must identify index HF hospitalizations, capture each one's discharge date, and count readmissions for HF admitted within 30 days of THAT discharge (per-patient, per-index linkage). Counting patients with multiple HF stays regardless of timing is the error.
| MS2-comp proportion · multi-step/compositional | not run |
| B88 | Among patients who had an ischemic stroke, what fraction had a carotid imaging study followed by a carotid endarterectomy within 6 months? Report the fraction.Trap and rubricCompositional cohort with ordered sub-steps: denominator = ischemic-stroke patients; numerator = those with carotid imaging AND a subsequent endarterectomy within 6 months of that imaging. A one-shot query flattens to 'stroke + imaging + endarterectomy present' and loses the ordering and the per-patient 6-month linkage. Grading: Must anchor on stroke patients, then require carotid imaging followed by endarterectomy within 6 months of the imaging, linked per patient. Counting co-occurrence of the three without ordering/windowing is the error.
| MS3-comp proportion · multi-step/compositional | not run |
| B89 | How many patients newly started on dialysis had at least one nephrology visit in the year BEFORE dialysis initiation? Report the count.Trap and rubricCompositional cohort spanning a pre-index window: step 1 = identify dialysis initiators and their initiation date; step 2 = look back 365 days from THAT date for a nephrology visit. A one-shot query typically checks 'has dialysis AND has nephrology visit' anywhere in history, ignoring that the visit must precede initiation within a specific window. Grading: Must find each patient's dialysis-initiation date, then require a nephrology visit in the 365 days before that date (per-patient pre-index window). Any-time co-occurrence is the error.
| MS4-comp count · multi-step/compositional | not run |
| B90 | How many patients received chemotherapy? Report the count.Trap and rubricMulti-domain de-duplication: chemotherapy exposure appears across DIFFERENT OMOP domains — procedure_occurrence (administration) and drug_exposure (oral/infused agents). A one-shot query checks one domain (undercount) or unions domains but counts rows, double-counting patients present in several. Correct handling unions the domains and counts DISTINCT patients once. Grading: Must capture chemotherapy across procedure_occurrence and drug_exposure and count DISTINCT patients (union, deduplicated), not rows-per-domain or a single domain. Single-domain counting or double-counting across domains is the error.
| MSRC-multisource count · multi-source-dedup | not run |
| B91 | For each patient, determine the first date on which they had at least two, at least three, at least four, and at least five distinct recorded chronic conditions. Report how many patients reached each milestone.Trap and rubricCumulative multimorbidity milestone dates: each threshold's date is when the CUMULATIVE count of DISTINCT chronic conditions first reaches it — requiring ordering each patient's first-diagnosis dates per distinct condition and finding when the running distinct-count hits 2,3,4,5. A one-shot query counts total conditions ever (ignoring the running accumulation and its timing) or counts condition rows (not distinct conditions). Correct handling accumulates distinct-condition onset over time. Grading: Must order each patient's per-condition first-diagnosis dates, track the running count of DISTINCT chronic conditions, and record when it first reaches each threshold. Counting total/ever conditions without the cumulative-over-time distinct logic is the error.
| MULTI1-increasing table · multimorbidity-milestones | not run |
| B92 | Estimate the one-year incidence of gout among adults with at least one year of continuous enrollment. Report the incidence.Trap and rubricPrevalence vs incidence: a pre-existing (prevalent) gout diagnosis is not an incident case and must be excluded via a baseline washout. Grading: Any answer excluding prevalent cases passes (washout present). Including them fails (no washout). The washout axis determines pass/fail. The enrollment axis (baseline-only vs full 24mo) determines which acceptable answer — diagnostic but not penalized.
| N1 proportion | not run |
| B93 | Count patients for whom NSAID use plausibly contributed to upper GI bleed within one year of starting the NSAID. Exclude cases where GI bleed is diagnosed within 60 days of starting the NSAID AND the patient had prodromal GI symptoms (dyspepsia/epigastric pain) before starting the NSAID (the NSAID was likely prescribed for the early symptoms of the bleed - reverse causation).Trap and rubricProtopathic bias / reverse causation: when a patient has prodromal symptoms before drug start AND the outcome appears within 60 days, the drug was likely prescribed FOR the developing disease. These cases must be excluded as reverse causation. Grading: System must check: (1) prodromal symptom before NSAID start AND (2) outcome within 60d. Both must be present to exclude. Boundary cases (within 60d but NO prodrome) are genuine.
| N10 count | not run |
| B94 | What is the rate of statin-induced rhabdomyolysis in the population? Report the rate.Trap and rubricWrong denominator: the at-risk denominator is statin-exposed patients, not the whole population. Rhabdomyolysis 'induced by statins' can only occur in patients who took statins. Grading: The system must restrict the denominator to statin-exposed patients. Dividing by total population commits this error. We grade the numerator COUNT (10) to avoid penalizing rate-method choice (proportion vs person-time rate).
| N2a proportion | not run |
| B95 | What is the prevalence of diabetic retinopathy in the population? Report the prevalence.Trap and rubricWrong denominator / disease-specific complication: diabetic retinopathy can only occur in diabetic patients, so the denominator must be restricted to diabetics. A system that divides by total population commits this error. Grading: System must restrict denominator to diabetic patients. Numerator = patients with diabetic retinopathy codes among diabetics.
| N2b proportion | not run |
| B96 | What is the rate of contrast-induced nephropathy in the population? Report the rate.Trap and rubricWrong denominator: the at-risk denominator is patients who received IV iodinated contrast media (a transient procedural exposure), not the whole population. Contrast-induced nephropathy can only occur in patients exposed to contrast. Subtlety vs N2a/N2b: the exposure is a transient PROCEDURE, not a chronic condition or a drug, so it is harder to recognize as the denominator restriction. Grading: System must restrict the denominator to patients who received IV contrast (CT/angiography with contrast). Numerator = acute kidney injury within 48-72h post-procedure among the contrast-exposed. Dividing by total population commits this error.
| N2c proportion | not run |
| B97 | What is the breast cancer screening completion rate among women aged 50-74? Report the rate.Trap and rubricWrong / multi-layer denominator: the eligible denominator is NOT all women 50-74. Standard quality-measure specs exclude women with (a) a prior bilateral mastectomy (no tissue to screen) and (b) an existing active breast cancer diagnosis (already diagnosed, screening is moot). The system must apply MULTIPLE clinical exclusion layers, not just demographic filtering. Grading: System must exclude, from the denominator, women with a prior bilateral mastectomy AND women with an existing active breast cancer diagnosis, before computing the screening rate. Rate = screened / truly eligible. Applying only the age/sex filter inflates the denominator and understates the rate. (Only bilateral mastectomy excludes; unilateral does not.)
| N2d proportion | not run |
| B98 | Measure the cervical-cancer screening rate among adult women. Exclude women with a prior total hysterectomy from the denominator (no cervix → not eligible for screening).Trap and rubricIneligibility by prior anatomic status: women who had a total hysterectomy are ineligible for cervical cancer screening. Missing this exclusion understates the screening rate by inflating the denominator with women who cannot be screened. Grading: System must exclude, from the denominator, women with documented absence of the cervix before computing the screening rate. Evidence of cervix absence may come from a procedure (total/complete hysterectomy — a partial/supracervical/subtotal hysterectomy leaves the cervix and does NOT qualify) OR an anatomical-status diagnosis (acquired/congenital absence of cervix). Rate = screened / eligible. Missing the exclusion inflates the denominator. The pass/fail axis is whether cervix-absent women are excluded at all — not the specific evidence source used.
| N3 proportion | not run |
| B99 | Build a cohort of new users of proton pump inhibitors (PPIs). Require at least 365 days of continuous prior observation before the first fill to confirm no earlier use. Count the number of eligible new users.Trap and rubricNew-user look-back / baseline observability: to confirm a patient is a new PPI user, verify >=365 days of prior observation (observation_period before the first PPI drug_exposure) AND no prior PPI in that window. Patients with insufficient look-back have unknown prior status and must be excluded. Grading: System must require >=365d prior observation (observation_period start to first PPI drug_exposure) AND no prior PPI in that window. Patients with short look-back are not confirmable new users.
| N4 count | not run |
| B100 | Estimate the one-year incidence of new-onset type 2 diabetes after statin initiation. Only include patients observed for the full 365-day window (or who had the event within that window); do not count unobserved time as event-free.Trap and rubricFollow-up sufficiency / immature outcome window: patients with insufficient post-index follow-up (data cut before 365d) must be excluded from the denominator unless they had the event. Counting them as event-free dilutes the incidence (understates risk). Grading: System must require full 365-day follow-up OR event within window. Patients whose observation ends before 365d without an event are NOT event-free — they are unobserved and must be excluded from the denominator.
| N5 proportion | not run |
| B101 | Define each patient's index date as the first qualifying diagnosis of rheumatoid arthritis that occurs during active enrollment and is a confirmed (not rule-out or history-of) diagnosis, then count patients who had a corticosteroid prescription in the 180 days before index.Trap and rubricIndex/anchor-date misspecification: the index must anchor on the first CONFIRMED, in-observation rheumatoid-arthritis condition_occurrence — not MIN(date) over all rows (which picks up 'Preliminary diagnosis'/rule-out records or pre-observation dates). A wrong anchor shifts the 180-day baseline window. Grading: System must anchor on the first RA condition_occurrence with a confirmed condition_status_concept_id (excluding 'Preliminary diagnosis'/'History of') AND condition_start_date within observation_period. Anchoring on MIN(date) of any RA code shifts the 180d window and over-counts corticosteroid use.
| N6 count | not run |
| B102 | Among patients in the cohort, report the proportion WITHOUT comorbidity Z at baseline. Only classify a patient as comorbidity-negative if they were observable at baseline (sufficient enrollment); do not infer 'negative' from patients with no baseline observation.Trap and rubricAbsence of record != absence of condition: treating patients with no baseline observation as comorbidity-negative overstates the healthy/negative fraction. Only observable patients can be classified negative. Grading: System must restrict the negative classification to patients observable at baseline (sufficient enrollment); unobserved patients are unknown and excluded, not counted as negative. Treating no-code as negative for everyone commits this error.
| N7 proportion | not run |
| B103 | Count patients with severe hyperglycemia (blood glucose above 250 mg/dL) at any measurement.Trap and rubricUnit harmonization: glucose is recorded in mixed units (mg/dL and mmol/L). A raw threshold on the number (>250) misclassifies mmol/L records. 250 mg/dL ≈ 13.9 mmol/L. Also: biologically implausible values should be filtered. Grading: System must convert mmol/L to mg/dL (×18.0182) before applying the 250 threshold, and filter biologically implausible values. Raw numeric comparison without unit checking commits this error.
| N8 count | not run |
| B104 | Build a cohort of new (incident) users of a statin, and report its size.Trap and rubricClass-wide new-user washout: a patient switching between statins (e.g. simvastatin -> atorvastatin) looks drug-incident but is class-prevalent. The washout must span the whole statin CLASS, assembled via the RxNorm concept hierarchy (concept_ancestor), not a single ingredient. Grading: Must apply a class-wide washout over all statins (no prior statin drug_exposure of any kind in the lookback), assembling the class via the RxNorm hierarchy, when defining incident users. Single-ingredient washout or no washout commits this error.
| ND1-classwashout count · drug-class/new-user | not run |
| B105 | What fraction of patients on antihypertensive therapy are adherent (proportion of days covered >= 0.8) in the year after initiation? Report the fraction.Trap and rubricClass-level adherence with overlap handling: PDC for 'antihypertensive therapy' spans a broad drug class, so the antihypertensive ingredients must be assembled via the RxNorm concept hierarchy (concept_ancestor). Overlapping drug_exposure intervals must be shifted forward (not double-counted) and inpatient days credited. A naive fill-count or double-counted-overlap adherence is the error. Grading: Must assemble the antihypertensive class via the RxNorm/ATC hierarchy, build a covered-days timeline from drug_exposure intervals shifting overlaps forward (and crediting inpatient days), compute PDC over 365 days, and count PDC>=0.8. Naive exposure-count adherence or double-counting overlaps is the error.
| ND2-pdc proportion · drug-class/adherence | not run |
| B106 | Among new users of an SSRI, what proportion switched to an SNRI (rather than adding one) within 12 months? Report the proportion.Trap and rubricSwitch-vs-augmentation across two drug CLASSES: a switch = SSRI exposure stops when SNRI starts; augmentation = both continue concurrently. Distinguishing them needs a 30-day overlap rule over drug_exposure intervals and assembly of BOTH the SSRI and SNRI classes via the RxNorm concept hierarchy. A one-shot query cannot express switch vs add. Grading: Must assemble both SSRI and SNRI classes (RxNorm hierarchy) and classify as a SWITCH only when the SSRI is discontinued around the SNRI start (<=30-day overlap of drug_exposure intervals), not augmentation. Counting anyone on both, or ignoring the overlap rule, is the error.
| ND3-classswitch proportion · drug-class/switching | not run |
| B107 | What is the rate of statin-associated myopathy in the population? Report the rate.Trap and rubricDrug-class exposure denominator: statin-associated myopathy can only occur in statin users, so the at-risk denominator is statin-exposed patients (the whole class, assembled via the RxNorm concept hierarchy), not the whole population. Dividing by the total population commits this error. Grading: Must restrict the denominator to statin-exposed patients (class-wide via the RxNorm hierarchy) and count myopathy among them. Dividing by the total population commits this error.
| ND4-classdenominator proportion · drug-class/denominator | not run |
| B108 | How many patients received pembrolizumab as their first line of therapy? Report the count.Trap and rubricLine-of-therapy construction: 'first line' requires building regimens from drug administrations, collapsing co-initiated agents (within ~28 days) into one line, ordering lines, and taking line 1 — not merely 'ever received pembrolizumab'. Even for a single agent, correctly attributing it to 1L requires LoT logic. (Marked drug-class/LoT; regimen partners may span classes -> safest to run post-fix.) Grading: Must construct lines of therapy (collapse co-initiated agents into a regimen, order lines) and count patients whose FIRST line contains pembrolizumab. Counting anyone ever exposed to pembrolizumab, ignoring line assignment, is the error.
| ND5-firstline count · drug-class/line-of-therapy | not run |
| B109 | How many older adults are on concurrent therapy with three or more anticholinergic medications? Report the count.Trap and rubricClass-level concurrency / polypharmacy: counts patients with >=3 DISTINCT anticholinergic ingredients whose drug_exposure intervals OVERLAP in time. Requires assembling the anticholinergic class via the RxNorm concept hierarchy and computing temporal overlap of >=3 distinct exposures. A one-shot query counts anyone with >=3 anticholinergic exposures ever, ignoring concurrency. Grading: Must assemble the anticholinergic class (RxNorm hierarchy) and require >=3 DISTINCT ingredients with concurrently overlapping drug_exposure intervals among older adults. Counting >=3 exposures anytime (no overlap, or same ingredient) is the error.
| ND6-concurrent count · drug-class/polypharmacy | not run |
| B110 | Among patients with hypertension, how many have never had an abnormal potassium result? Report the count, and separately report how many have no potassium measurement at all.Trap and rubricNever-abnormal vs never-measured: 'never had an abnormal potassium' is ambiguous with the absence of testing — a patient with ZERO potassium labs is not the same as one tested and always normal. A one-shot query lumps never-measured patients into 'never abnormal', silently treating missing data as normal and inflating the count. Correct handling separates three groups: tested-and-always-normal, tested-with-an-abnormal, and never-tested (reported as its own bucket). Grading: Must distinguish never-abnormal-BUT-tested from never-MEASURED (zero potassium labs), reporting never-tested as a separate bucket rather than counting them as 'never abnormal'. Treating absence of testing as a normal result is the error.
| NEG1-nevermeasured table · negation-missingness | not run |
| B111 | Among patients with newly diagnosed epilepsy who started an antiepileptic drug, what percentage achieved seizure freedom (no recorded seizure) within one year of starting therapy? Report the percentage.Trap and rubricAbsence-of-event-in-window as a positive outcome: 'seizure freedom' = NO seizure code in the 1-year window AFTER AED initiation, among patients OBSERVABLE for that full year. This is temporal negation over an observable window. A one-shot query tends to count patients with a seizure (the opposite), or treats patients lost to follow-up as event-free, or ignores the post-initiation window. Correct handling requires full-year observability and the ABSENCE of a seizure in that window. Grading: Must define the numerator as patients with NO seizure code in the 365 days after AED start, restricted to patients observable for that full year (not lost to follow-up). Counting seizures, or treating unobserved patients as seizure-free, is the error.
| NG1-eventfree proportion · absence-in-window | not run |
| B112 | What is the median duration of chronic-condition treatment episodes, where episodes still ongoing at the end of the data have no recorded end date? Report the median in days.Trap and rubricOpen-ended interval handling: episodes with a NULL/absent end date are still ONGOING, not zero-length or erroneous. Duration math on a NULL end silently drops those rows (NULL arithmetic) or treats them as instantaneous, biasing the median toward shorter, completed episodes. Correct handling censors open episodes at the data cut-off (or last activity date) and includes them, acknowledging they are minimum durations. A one-shot query computes end-minus-start, losing every ongoing episode. Grading: Must impute an end for open (NULL-end) episodes at the data cut-off / last-activity date and include them (as censored/minimum durations), not drop them or treat them as zero. Letting NULL-end arithmetic silently exclude ongoing episodes is the error.
| NULLEND1-openinterval summary-statistic · open-interval-censoring | not run |
| B113 | What fraction of patients received a statin within 90 days of their first ASCVD diagnosis? Report the fraction.Trap and rubricExposure-opportunity denominator: the denominator must be restricted to patients ENROLLED for the full 90-day opportunity window after first ASCVD diagnosis — a patient who disenrolls on day 20 had no chance to be observed filling a statin and biases the fraction downward if kept. A one-shot query uses all ASCVD patients as the denominator regardless of whether they were observable for 90 days. Grading: Must restrict the denominator to first-ASCVD patients enrolled for the full 90-day post-diagnosis window (observable), then compute the fraction with a statin fill in that window. Using all ASCVD patients regardless of follow-up observability is the error.
| OPP1-opportunity proportion · exposure-opportunity-denominator | not run |
| B114 | How many patients alternated between controlled and uncontrolled diabetes (by HbA1c) at least three times? Report the count.Trap and rubricState oscillation detection: 'alternated at least three times' means the ordered state sequence has >=3 transitions between the two states — requiring ordered per-patient state records and counting actual switches, not merely presence of both states. A one-shot query counts patients who have both a controlled and an uncontrolled result (which needs only one of each), vastly over-counting. Correct handling orders states over time and counts transitions. Grading: Must order each patient's control-state records over time and count patients with >=3 transitions between states. Counting patients who merely have both states present is the error.
| OSC1-alternating count · state-oscillation | not run |
| B115 | How many patients on warfarin had a gastrointestinal bleed while also taking an NSAID? Report the count.Trap and rubricEvent-during-overlap window: the GI bleed must occur DURING concurrent warfarin+NSAID exposure (>=14 overlapping supplied days from drug_exposure intervals), not merely anytime in a patient who ever took both. A one-shot query checks 'has warfarin AND NSAID AND GI bleed' anywhere, ignoring that the bleed must fall inside the co-exposure overlap. Correct handling computes the overlap interval from drug_exposure_start_date/end_date and requires the bleed within it. Grading: Must compute the concurrent warfarin+NSAID overlap window (intersection of drug_exposure intervals) and count GI bleeds (condition_occurrence) occurring WITHIN that overlap, per patient. Counting patients with all three anytime, ignoring the overlap window, is the error.
| OV1-overlap count · co-exposure-window | not run |
| B116 | Build a cohort of patients who are new (incident) recipients of chronic hemodialysis, and report its size.Trap and rubricBaseline observability / new-user look-back, procedure-anchored: to call a dialysis recipient 'new', the system must confirm no prior dialysis AND require adequate prior observation (>=365 days of enrollment before the first dialysis). Patients with a first OBSERVED dialysis but insufficient look-back may be prevalent recipients whose earlier procedures are simply unobserved (left-truncation). The look-back requirement is HIDDEN — the system must discover it. Grading: System should require >=365 days of continuous prior observation (observation_period) before the first observed dialysis procedure_occurrence AND no prior dialysis in that window, before counting a patient as a new recipient. Treating any first observed dialysis as 'new' (ignoring look-back) commits this error.
| P1-newuser count | not run |
| B117 | Estimate the one-year incidence of new-onset heart failure after coronary artery bypass grafting (CABG). Report the incidence.Trap and rubricFollow-up sufficiency / immature outcome window, procedure-anchored: patients whose observation ends before 365 days post-CABG without a heart-failure diagnosis are NOT event-free — they are unobserved and must be excluded from the denominator (or handled by censoring). Counting their unobserved time as 'no event' dilutes the incidence. The follow-up requirement is HIDDEN. Grading: System should require a full 365-day post-CABG follow-up window (or an event within it), excluding patients censored before 365d with no event, or use a censoring-aware estimator. Using all CABG patients as the denominator and treating unobserved time as event-free commits this error.
| P2-followup proportion | not run |
| B118 | Type 2 diabetes can be identified by a qualifying diagnosis, a qualifying lab (HbA1c>=6.5), or a glucose-lowering medication. For patients who qualify, report which pathway qualified them FIRST (earliest qualifying date), and break down counts by that first pathway.Trap and rubricCompeting-pathway provenance: multiple independent definitions can each make a patient a case; the answer needs, per patient, the EARLIEST qualifying date across pathways and WHICH pathway that was — then a breakdown by first-qualifying pathway. A one-shot query picks one pathway, or reports overlapping per-pathway totals, and cannot attribute the earliest-qualifying source. Correct handling computes each pathway's first date per patient and takes the argmin. Grading: Must compute, per patient, the earliest qualifying date under each of the three pathways, pick the minimum (the qualifying pathway), and break down patient counts by that first pathway (ties handled deterministically). Reporting per-pathway totals or a single pathway is the error.
| PATH1-provenance table · pathway-provenance | not run |
| B119 | What is the median number of outpatient visits per patient in 2022? Report the median.Trap and rubricStatistic over the correct unit of analysis: 'median visits per patient' requires FIRST computing each patient's visit count, THEN taking the median across patients. A one-shot query often takes a median/percentile over the visit rows directly, or divides total visits by patients (a mean, not a median), giving a different number. Correct handling aggregates to one value per patient before computing the percentile. Grading: Must compute per-patient visit counts and then the median across patients (one value per patient). Taking a percentile over raw visit rows, or reporting mean visits/patient as the 'median', is the error.
| PCTL1-perpatient summary-statistic · per-unit-statistic | not run |
| B120 | How many patients progressed through at least two successively higher chronic kidney disease stages (without an intervening lower stage)? Report the count.Trap and rubricMonotonic ordered progression (state-machine over records): must order each patient's CKD stage records over time and detect a strictly non-decreasing advance through >=2 higher stages with NO intervening lower-stage record. A one-shot query typically checks 'has stage 3 and stage 4 anytime', ignoring order and intervening reversals. Correct handling reasons over the ordered stage sequence per patient. Grading: Must order stage records per patient and require >=2 successive upward stage transitions with no intervening lower stage. Checking co-occurrence of two stages regardless of order/reversal is the error.
| PRG1-progression count · ordered-progression | not run |
| B121 | Assign each patient in 2022 to the provider responsible for the greatest number of their qualifying outpatient visits, resolving ties by the most recent visit then the lowest provider identifier. Report the number of patients attributed to each provider's specialty.Trap and rubricPlurality provider attribution with deterministic ties: each patient maps to ONE provider — the plurality-visit provider — with a stated tie-break, so a patient is counted once. A one-shot query counts patient-provider pairs (a patient seeing three providers appears three times) or picks an arbitrary provider on ties. Correct handling ranks providers by visit count per patient, applies the tie-break, and attributes each patient once. Grading: Must attribute each patient to their single plurality-visit provider (tie-break: most recent visit, then lowest provider id), each patient counted once. Counting patient-provider pairs or arbitrary tie handling is the error.
| PROV1-plurality table · provider-attribution | not run |
| B122 | What is the point prevalence of COPD on January 1, 2023? Report the prevalence.Trap and rubricPoint-in-time denominator: point prevalence on a specific date counts, in the numerator, patients with a COPD diagnosis in a defined lookback (e.g. prior 3 years) who are ENROLLED on that exact date; the denominator is patients enrolled on that exact date. A one-shot query typically uses everyone with a COPD code ever / all patients, ignoring the on-date enrollment requirement for both numerator and denominator. Grading: Must restrict both numerator and denominator to patients enrolled on 2023-01-01 (enrollment span covering that date), with the numerator having a COPD diagnosis in the prior window. Using all-time codes / all patients, ignoring on-date enrollment, is the error.
| PT1-pointprev proportion · point-in-time-denominator | not run |
| B123 | What is the prevalence of asthma in 2022? Report the prevalence.Trap and rubricPrevalence-definition ambiguity: 'prevalence in 2022' can mean PERIOD prevalence (any asthma diagnosis during 2022 among those enrolled in 2022) or POINT/annual prevalence with a lookback (asthma dx in a prior window while enrolled). These give materially different counts. A defensible system states which definition it uses and applies matching enrollment; a one-shot query silently picks 'any code in 2022 / all patients' without acknowledging the choice or aligning the denominator. Grading: Must adopt a coherent prevalence definition (period or point) and align the denominator to the enrolled population for that definition, stating the choice. Silently counting any asthma code in 2022 over all patients (mismatched denominator, undeclared definition) is the error.
| PT2-periodvspoint proportion · interpretation-divergence | not run |
| B124 | What is the incidence of gout per 1,000 person-years over 2020-2022? Report the incidence.Trap and rubricMulti-span person-time: patients often have MULTIPLE enrollment spans (disenroll then re-enroll). Person-years at risk must be SUMMED across all qualifying spans within 2020-2022 (and stop at the first gout event), not taken from only the latest/longest span. A one-shot query typically uses a single span or a naive max-min date range, mis-stating the denominator person-time. Grading: Must sum at-risk person-time across each patient's multiple enrollment spans within the window (censoring at first gout, disenrollment, or window end), and count incident (first) gout. Using a single span or last-minus-first dates for person-time is the error.
| PT3-multispan rate · person-time-multispan | not run |
| B125 | Calculate outpatient-observable person-time for 2022, excluding days each patient spent as an inpatient. Report total person-years.Trap and rubricPerson-time with carved-out intervals: Inpatient Visit stays must be SUBTRACTED from each patient's observable interval (observation_period intersected with 2022), which requires interval difference against the union of inpatient visit_occurrence intervals, handling overlapping/adjacent stays so no inpatient day is double-subtracted. A one-shot query subtracts a raw count of inpatient rows/days or ignores inpatient time. Grading: Must subtract the MERGED Inpatient Visit intervals (visit_start_date/visit_end_date) from each patient's observable interval (interval difference), not a raw inpatient-day count, then sum remaining person-time. Subtracting raw inpatient rows (overlap double-count) or ignoring inpatient time is the error.
| PTINP1-excludeinpatient summary-statistic · person-time-exclusion | not run |
| B126 | Split each patient's observable time across calendar month, calendar year, age group, and diabetes disease-status (pre- vs post-diagnosis) boundaries, and report total person-days by each combination without double-counting any day.Trap and rubricMultidimensional person-time splitting: each observable interval must be cut simultaneously at month-ends, year-ends, birthdays, AND the diagnosis date, so every day belongs to exactly ONE (month, year, age, disease-status) cell with no day counted twice or dropped. A one-shot query splits on at most one dimension, or assigns whole intervals to a single cell, producing overlapping/duplicated or missing person-days. Correct handling intersects all boundary sets and verifies the day total reconciles. Grading: Must split intervals at the union of month, year, birthday, and diagnosis boundaries so each day maps to exactly one multidimensional cell (total person-days reconciles to raw observable days). Splitting on one dimension or assigning whole intervals is the error.
| PTMD1-multidim table · person-time-multidim | not run |
| B127 | Compute each patient's total observable person-days two ways — by expanding to one row per observable day and by summing merged date intervals — and report whether the two totals agree.Trap and rubricPerson-time method reconciliation: the daily-expansion total and the merged-interval total must AGREE only if overlapping enrollment spans are correctly merged and interval endpoints are counted consistently (inclusive/exclusive). Discrepancies reveal double-counted overlap days or off-by-one endpoint errors. A one-shot query computes one method (often raw interval sums that double-count overlaps) and never cross-checks. Correct handling implements both and reconciles, exposing overlap/boundary handling. Grading: Must compute person-days by day-level expansion AND by summing merged intervals, then compare — agreement requires merging overlaps and consistent endpoint counting. Reporting a single method (esp. raw interval sums that double-count overlaps) with no reconciliation is the error.
| PTREC1-reconcile table · person-time-reconciliation | not run |
| B128 | Report total observed person-years by calendar year and age group, correctly splitting each patient's observable time when they cross a year boundary or have a birthday mid-interval. Report the grid.Trap and rubricPerson-time splitting across year AND age boundaries: a single observable interval must be CUT at each Dec 31 and at each birthday, allocating the right fraction of person-time to each (year, age-group) cell. A one-shot query assigns a patient's whole interval to one year/age (by index date or enrollment start), mis-allocating person-time. Correct handling splits intervals at both boundary types. Grading: Must split each observable interval at calendar-year boundaries and at birthdays, attributing the correct person-time fraction to each (year, age-group) cell. Assigning the whole interval to a single year/age bucket is the error.
| PY1-yearage table · person-time-split | not run |
| B129 | Report both the annual and the monthly counts of active heart-failure episodes for 2022, and reconcile them — an episode spanning multiple months must not be summed across months to equal the annual total. Report both plus the reconciliation.Trap and rubricCross-grain reconciliation: an episode active across several months is counted in EACH of those months, so summing monthly counts overstates the (distinct-episode) annual total. The system must compute annual as distinct episodes active in the year, monthly as episodes active per month, and explain the difference (episodes spanning months). A one-shot query reports one grain, or naively sums months to 'annual', producing an inflated, inconsistent total. Grading: Must compute the annual count as distinct episodes active in 2022 AND monthly active-episode counts, recognizing that monthly counts sum to MORE than annual because multi-month episodes recur; the totals are reconciled, not equated. Summing months to get the annual total is the error.
| REC1-reconcile table · aggregation-reconciliation | not run |
| B130 | How many acute myocardial infarction events occurred in 2022, where events for the same patient within 30 days of each other count as a single event? Report the event count.Trap and rubricRecurrent-event counting with a refractory window: the same clinical event generates repeat codes across days; a NEW event only counts if it is >30 days after the prior counted event for that patient. This differs from first-event-only (undercount) and from raw code counts (overcount). It also differs from episode-collapsing by adjacency — here re-qualification requires a fixed clearance gap. Correct handling walks each patient's ordered event dates, starting a new event only when >30 days have elapsed since the last counted one. Grading: Must count recurrent events per patient where a new event requires >30 days since the previously counted event (collapsing codes within 30 days into one), allowing multiple events per patient. Counting all codes (overcount) or only the first event (undercount) is the error.
| RECUR1-refractory count · recurrent-event | not run |
| B131 | How many patients qualified for a heart-failure cohort, stopped qualifying, and then re-qualified at least 180 days after they last stopped qualifying? Report the count.Trap and rubricRequalification after a clearance gap: membership is a time-varying interval; requalification requires the patient to LEAVE the cohort and re-enter only after >=180 days of non-qualifying time. A one-shot query treats cohort membership as a static ever/never flag and cannot express leave-then-return, or counts anyone with two qualifying spans regardless of the 180-day gap. Correct handling builds per-patient qualifying intervals and detects a re-entry preceded by a >=180-day non-qualifying gap. Grading: Must construct time-varying cohort-membership intervals per patient and count those with a re-entry occurring >=180 days after a prior exit. A static ever-qualified flag, or ignoring the 180-day gap, is the error.
| REQUAL1-requalify count · cohort-requalification | not run |
| B132 | For a single named maintenance medication (levothyroxine), reconstruct exposure when refills arrive before the prior supply is exhausted, carrying unused supply forward. Report the median continuous exposure duration.Trap and rubricRefill stockpiling / carry-forward: early refills mean supplied days accumulate ahead of consumption; a correct exposure timeline shifts each levothyroxine drug_exposure's start to the end of the prior adjusted interval (carry-forward), rather than starting at the raw drug_exposure_start_date. A one-shot query sums raw interval lengths (double-counting overlap) or uses fill-date-to-fill-date spacing. Grading: Must build the exposure timeline by carrying forward unused supply (each exposure's coverage begins when the prior adjusted interval ends), then measure continuous exposure. Summing raw drug_exposure interval lengths over overlaps, or raw start-date spacing, is the error.
| RX1-stockpile summary-statistic · refill-stockpiling | not run |
| B133 | Report the distribution of the Charlson Comorbidity Index among patients hospitalized for pneumonia. Report the distribution.Trap and rubricComposite clinical score from many components: the Charlson index is computed per patient by mapping ~17 weighted comorbidity categories (MI, CHF, dementia, diabetes, cancer, liver disease, etc.) from diagnosis codes and summing the weights. A one-shot query cannot assemble a 17-category weighted score in a single pass — it tends to count raw diagnoses or a single condition. Correct handling builds each comorbidity flag per patient, applies the Charlson weights, sums per patient, then reports the score distribution. Grading: Must map the ~17 Charlson comorbidity categories from diagnosis codes per patient, apply the standard category weights, sum to a per-patient score, and report the distribution. Counting raw diagnoses or omitting the weighted multi-category composition is the error.
| SC1-charlson table · composite-score | not run |
| B134 | Is there a seasonal pattern in new influenza diagnoses? Report the diagnosis rate by calendar month, pooled across 2019-2023.Trap and rubricSeasonality with person-time / days-in-month normalization: raw monthly COUNTS conflate true seasonality with (a) unequal days per month (Feb vs Jul) and (b) changing enrollment/observable population by month/year. A one-shot query reports raw counts per month, implying a seasonal signal that is partly a denominator artifact. Correct handling divides monthly events by observable person-time (or population) that month, and pools rates comparably across months. Grading: Must express each month as a RATE normalized by observable person-time/population that month (accounting for days-in-month and enrollment), not raw counts, before comparing months. Reporting raw monthly counts as the seasonal signal is the error.
| SEAS1-monthnorm table · seasonality-normalization | not run |
| B135 | Convert each patient's overlapping active chronic-condition periods into non-overlapping time segments, each labeled by the exact set of conditions active during that segment. Report the segment breakdown.Trap and rubricInterval flattening into labeled segments: overlapping condition intervals must be cut at every start/stop boundary into disjoint segments, each tagged with the SET of conditions simultaneously active. A one-shot query cannot decompose overlapping intervals into a segment timeline; it reports per-condition durations that overlap and double-count time. Correct handling is a sweep-line / boundary-split over interval unions. Grading: Must split the union of overlapping condition intervals at every boundary into non-overlapping segments, each labeled by the active condition set, so total segment time has no double-counting. Reporting overlapping per-condition intervals is the error.
| SEG1-flatten table · interval-flatten | not run |
| B136 | Count patients who had, in order, a qualifying diagnosis, then a confirmatory test, then a treatment initiation, all within a single 90-day window (unrelated events may occur in between). Report the count.Trap and rubricOrdered subsequence within a window: the three events must occur in the specified ORDER and within 90 days of each other, but other unrelated events may interleave. A one-shot query checks co-occurrence of the three within 90 days without enforcing order, or requires them to be strictly consecutive records. Correct handling finds an ordered subsequence (dx < test < treatment) inside a 90-day span per patient. Grading: Must require dx-date < test-date < treatment-date with the span from dx to treatment <=90 days, per patient, allowing unrelated intervening events. Ignoring order, or requiring strict adjacency, is the error.
| SEQ1-subsequence count · ordered-subsequence | not run |
| B137 | How many patients completed a full 3-dose hepatitis B vaccination series, respecting the minimum required intervals between doses? Report the count.Trap and rubricDose-series construction with spacing rules: a valid series requires 3 doses in order with MINIMUM intervals between them (e.g. dose2 >= 4 weeks after dose1, dose3 >= 8 weeks after dose2 and >= 16 after dose1); doses too close don't count and same-day duplicates collapse. A one-shot query counts patients with >=3 vaccine records, ignoring spacing and duplicates. Correct handling walks the ordered doses applying the interval rules. Grading: Must validate an ordered 3-dose series meeting the minimum inter-dose intervals (collapsing same-day duplicates, rejecting too-close doses). Counting patients with >=3 vaccine records regardless of spacing is the error.
| SER1-doseseries count · series-construction | not run |
| B138 | Using the code-based, lab-based, and medication-based definitions of diabetes, report how many patients meet exactly one, exactly two, or all three definitions. Report the mutually exclusive counts.Trap and rubricMutually-exclusive set partitioning (no double-count): a patient meeting 2 diabetes definitions (condition_occurrence codes, measurement-based lab thresholds, and drug_exposure) must be counted once in the 'exactly two' bucket, not in each. Requires per-patient membership across the three definitions and partitioning by the COUNT of definitions met. A one-shot query reports each definition's total (overlapping) or a simple union, which double-counts. Grading: Must compute, per patient, how many of the three diabetes definitions are met, then bucket patients by that count (exactly 1 / exactly 2 / all 3) with each patient in exactly one bucket. Reporting overlapping per-definition totals is the error.
| SET1-exclusive table · set-partitioning | not run |
| B139 | Among patients with any malignancy diagnosis, how many have non-melanoma skin cancer as their ONLY malignancy? Report the count.Trap and rubricPatient-level set subtraction, not code-level: 'only malignancy is NMSC' means the patient has >=1 NMSC code AND ZERO codes for any other malignancy — a per-PATIENT condition over their full code set, not a per-code filter. A one-shot query filters to NMSC codes (keeping patients who also have other cancers) or subtracts at the code level, both wrong. Correct handling computes, per patient, the set of distinct malignancy types and keeps those whose set == {NMSC}. Grading: Must identify patients whose ENTIRE set of malignancy diagnoses contains only non-melanoma skin cancer (>=1 NMSC and no other malignancy code), a patient-level set condition. Filtering to NMSC rows (ignoring patients' other cancers) or code-level subtraction is the error.
| SET2-onlysubset count · set-subtraction | not run |
| B140 | How many patients have chronic kidney disease stage 3 or worse? Report the count.Trap and rubricLab-based staging with repeated-measure confirmation: CKD stage 3+ is defined physiologically by TWO eGFR measurements < 60 taken at least 90 days apart (to establish chronicity), not by a single low eGFR and not by CKD stage billing codes (under-coded). A one-shot query uses stage codes, or a single eGFR<60 (which may be acute kidney injury, not chronic). Correct handling requires two qualifying eGFR values >=90 days apart per patient. Grading: Must define CKD 3+ from >=2 eGFR measurements <60 at least 90 days apart per patient (chronicity), computing eGFR from creatinine if needed — not from CKD stage codes and not from a single low eGFR. Single-measurement or code-based definitions commit this error.
| ST1-eGFRstage count · lab-based-staging | not run |
| B141 | How many patients had at least three consecutive abnormal laboratory results with no normal result in between? Report the count.Trap and rubricConsecutive-run / streak detection: must order each patient's results by date and find a run of >=3 abnormal results uninterrupted by any normal result. A one-shot query counts patients with >=3 abnormal results total (ignoring that a normal result in between breaks the streak). Correct handling is gap-and-island/streak logic over the ordered sequence. Grading: Must detect a maximal run of >=3 consecutive abnormal results (ordered by date) with no intervening normal result per patient. Counting >=3 abnormal results anywhere (streak-agnostic) is the error.
| STK1-streak count · consecutive-run | not run |
| B142 | Report the prevalence of obesity by calendar year and sex for 2020-2022. Some patients have sex recorded inconsistently across years.Trap and rubricInconsistent time-varying attribute across strata: when a patient's recorded sex differs across years, naive stratification places the same patient in different sex strata in different years (or double-counts), producing incoherent sex-specific denominators. A one-shot query strata by the per-row/per-year value without reconciling. Correct handling resolves each patient to ONE sex (e.g. most-recent or modal, stated) and applies it consistently, or explicitly reports the inconsistency count. Grading: Must resolve each patient's sex to a single value (stated rule, e.g. modal/most-recent) applied consistently across years, or flag inconsistent patients as a separate bucket — not let the same patient flip strata year to year. Stratifying by unreconciled per-year sex is the error.
| STR1-attrconsistency table · strata-consistency | not run |
| B143 | How many patients had four or more emergency-department visits within any rolling 90-day period? Report the count.Trap and rubricSliding-window maximum (not fixed calendar windows): 'any rolling 90-day period' means the window can start on ANY visit date, so a patient with 4 visits spanning day 10-95 qualifies even though no single calendar quarter contains all four. A one-shot query buckets by fixed calendar quarter/year and misses cross-boundary clusters. Correct handling checks, per patient, whether any visit-anchored 90-day window contains >=4 visits. Grading: Must evaluate rolling 90-day windows anchored at each ED visit per patient (e.g. count visits within 90 days of each visit) and flag patients reaching >=4 in any such window. Fixed calendar-period bucketing is the error.
| SW1-slidingmax count · sliding-window-max | not run |
| B144 | Classify each patient's terminal state as death, end of observability (disenrollment), or end of available data, and report the distribution.Trap and rubricTerminal-state classification: a patient's follow-up ends for one of three DISTINCT reasons — death, disenrollment, or the data cut-off — and conflating them (e.g. treating disenrollment or data-end as death, or all non-deaths as 'still followed') misrepresents censoring. A one-shot query typically has no notion of why observation stops. Correct handling assigns each patient exactly one terminal reason using the earliest applicable of death date, disenrollment date, and global data-end. Grading: Must classify each patient's follow-up end as death vs disenrollment vs data-end (mutually exclusive, earliest applicable), not conflate them. Ignoring the reason observation stops, or equating disenrollment/data-end with death, is the error.
| TERM1-terminalstate table · terminal-state-classification | not run |
| B145 | For each patient, identify their first qualifying diabetes diagnosis and report the count by diagnosis type; when multiple qualifying diagnoses occur on the same first date, count the patient once.Trap and rubricSame-day tie-break for 'first' event: when a patient has multiple qualifying diagnoses on their earliest date, 'the first' is ambiguous — a naive MIN(date) join returns MULTIPLE rows for that patient, double-counting them across diagnosis types. A one-shot query joins on the minimum date without collapsing ties, inflating per-type counts and total > patient count. Correct handling applies a deterministic tie-break (or counts the patient once) so each patient contributes exactly one first event. Grading: Must ensure each patient contributes exactly one first event despite same-date ties (deterministic tie-break or count-once), so per-type counts sum to the distinct patient count. A MIN(date) join that returns multiple same-day rows per patient is the error.
| TIE1-samedaytie table · first-event-tiebreak | not run |
| B146 | Produce a transition matrix of counts of observed movements between consecutive recorded CKD stages across all patients (e.g. stage 2 to stage 3), collapsing consecutive same-stage records first. Report the matrix.Trap and rubricConsecutive-state transition counting: a transition is a change between a patient's TEMPORALLY ADJACENT distinct states after collapsing repeats — so stage2,stage2,stage3 yields one 2->3 transition, not two. A one-shot query cross-joins all stage pairs per patient (counting non-adjacent and self pairs) or counts every record pair, inflating the matrix. Correct handling orders states, collapses runs, and tallies adjacent (from,to) pairs. Grading: Must collapse consecutive same-stage records, then count only temporally adjacent (from->to) stage pairs into the matrix. Cross-joining all stage pairs or counting non-collapsed adjacent duplicates is the error.
| TRANS1-matrix table · transition-matrix | not run |
| B147 | Count patients who had an abnormal cardiac stress test that was followed by a coronary angiography within 90 days, and then a revascularization procedure. Report the count.Trap and rubricEvent-ordering dependency: the answer requires a strict temporal SEQUENCE — abnormal stress test, THEN angiography within 90 days, THEN revascularization (after the angiography). A one-shot query typically writes flat co-occurrence filters and ignores ordering, counting patients who had all three in any order. Correct handling reasons about per-patient event order and inter-event windows. Grading: System must enforce the temporal order (stress test -> angiography within 90d -> later revascularization) using event dates per patient. Counting patients who have all three procedures without ordering/windowing is the error.
| TS1-order count · temporal-sequence | not run |
| B148 | Count patients whose first-ever diagnosis of chronic kidney disease occurred AFTER their first diagnosis of type 2 diabetes. Report the count.Trap and rubricFirst-event ordering: the answer depends on comparing the FIRST occurrence date of two conditions per patient (CKD onset after T2D onset). A one-shot query that just requires both diagnoses present (or compares any CKD date to any T2D date) mis-answers. Correct handling computes MIN(date) per condition per patient and compares the anchors. Grading: System must compute each patient's first (earliest) T2D date and first CKD date and count only those where first CKD > first T2D. Using any-date co-occurrence, or not anchoring on first occurrence, is the error.
| TS2-order count · temporal-sequence | not run |
| B149 | Count patients whose first opioid prescription came AFTER their first documented chronic pain diagnosis (not before). Report the count.Trap and rubricFirst-event ordering across two domains: requires comparing MIN(chronic-pain condition_occurrence date) to MIN(opioid drug_exposure date) per patient and keeping only opioid-after-pain. A one-shot query that merely requires both present, or compares arbitrary dates, mis-answers. Grading: Must compute first pain-diagnosis date (condition_occurrence) and first opioid drug_exposure date per patient and count only those with first opioid > first pain. Co-occurrence or non-first-event date comparison is the error.
| TS3-order count · temporal-sequence | not run |
| B150 | Count patients diagnosed with lung cancer whose diagnosis was preceded by a chest CT within the prior 90 days (a workup-then-diagnosis pattern). Report the count.Trap and rubricDirectional pre-event window (ordering): the CT must come BEFORE the cancer diagnosis, within 90 days prior. A one-shot query commonly checks 'has lung cancer AND has chest CT' or uses a symmetric window, capturing post-diagnosis surveillance CTs too. Correct handling enforces CT-before-diagnosis within the prior-90-day window. Grading: Must require a chest CT dated within the 90 days BEFORE the lung-cancer diagnosis, per patient (directional). Counting any CT (including after diagnosis) or ignoring the window is the error.
| TS4-order count · temporal-sequence | not run |
| B151 | Among patients newly diagnosed with heart failure, what is the median time to first hospitalization? Report the median in days.Trap and rubricTime-to-event with right-censoring: patients who disenroll or reach data end WITHOUT being hospitalized are censored, not absent. Taking the median only over patients who were hospitalized (ignoring censored follow-up) badly underestimates the median time — and the true median may be un-reached (>50% never hospitalized). A one-shot query computes the mean/median of observed event times among those with the event. Correct handling accounts for censored follow-up (Kaplan-Meier median or explicit at-risk reasoning), or states the median is not reached. Grading: Must incorporate censored follow-up (patients without the event contribute at-risk time; median from a survival/at-risk estimate, or reported as not-reached if <50% have the event). Taking the median of event times only among those hospitalized is the error.
| TTE1-censored summary-statistic · time-to-event-censoring | not run |
| B152 | For heart failure in 2022, report the count at five levels: distinct patients, distinct hospitalization episodes, distinct encounters, distinct records, and records. Report all five.Trap and rubricUnit-of-analysis integrity: the same clinical reality yields very different numbers at patient / hospitalization-episode / visit / condition-record grain, and conflating them is a classic RWE error. The question forces each grain to be computed correctly (distinct persons, collapsed inpatient episodes, distinct visit_occurrence, raw condition_occurrence rows). A one-shot query returns one COUNT(*) at whatever grain the table happens to be. Grading: Must produce distinct counts at the patient, hospitalization-episode (collapsed Inpatient Visits), visit, and condition-record grains — each at its correct level of aggregation. A single COUNT at one grain is the error.
| UOA1-grid table · unit-of-analysis | not run |
| B153 | How many patients had a follow-up hepatitis B vaccination within one year of their first dose? Report the count under each interpretation of 'within one year': within 365 days, within 366 days on a leap year, and on or before the same calendar date next year.Trap and rubricCompeting definitions of 'within one year': '365 days', 'the same calendar date next year' (which is 366 days across a leap year), and 'by the first anniversary' are DIFFERENT boundaries yielding different counts for events near the edge. A one-shot query silently equates 'one year' with a single arithmetic (usually 365 days) and hides the divergence. Correct handling computes each interpretation and surfaces that the answer depends on the definition. Grading: Must compute the count under each distinct 'within one year' interpretation (365-day, calendar-anniversary incl. leap-year effect) and surface the divergence, not silently pick one. Collapsing all senses into a single 365-day rule is the error.
| WITHIN1-oneyear table · interpretation-divergence | not run |
Tasks that create objects in Linkr, which the user then reviews and edits in the interface. Database: MIMIC-IV demo in OMOP CDM. Expected results are produced by hand, independently of the runs. Three runs per task. Values in brackets are not fixed yet.
Prompt copies the instruction to paste into a new conversation of an MCP client connected to Linkr.
| # | Task | Expected output | Success | Result |
|---|---|---|---|---|
| C1 | Cohort: adults with an ICU stay of 48 hours or more | Patient list | Same list | not run |
| C2 | Cohort: vancomycin given in the ICU | Patient list | Same list | not run |
| C3 | Cohort: lactate above 4 mmol/L within 24 hours of ICU admission | Patient list | Same list | not run |
| C4 | Edit C1 to exclude deaths within 24 hours, report the attrition | Patient list, attrition | Both identical | not run |
| C5 | Import an OHDSI Phenotype Library definition and run it | Patient list | Same list | not run |
| C6 | Map 20 laboratory source codes | OHDSI mapping | [threshold] correct | not run |
| C7 | Map 20 drug source codes | OHDSI mapping | [threshold] correct | not run |
| C8 | Dataset from C1: age, sex, length of stay, first-day maximum lactate, documented | Table | Same values | not run |
| C9 | R script: descriptive table of C1 | Table | Same values | not run |
| C10 | Python script: distribution of the first lactate, saved figure | Quartiles, file | Same quartiles, file present | not run |
| C11 | Dashboard for C1: patient count, age histogram, sex distribution | 3 widgets | 3 of 3 | not run |
| # | Harness | Model | Answer | Verdict | Tools | Tokens in | Batch |
|---|
Click a row to see the prompt, the tool calls, the model's last reply and the verdict.