[
  {
    "title": "Accuracy",
    "header": [
      {
        "value": "Model",
        "markdown": false,
        "metadata": {}
      },
      {
        "value": "Mean win rate",
        "description": "How many models this model outperforms on average (over columns).",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {}
      },
      {
        "value": "MedCalc-Bench - MedCalc Accuracy",
        "description": "MedCalc-Bench is a benchmark designed to evaluate models on their ability to compute clinically relevant values from patient notes. Each instance consists of a clinical note describing the patient's condition, a diagnostic question targeting a specific medical value, and a ground truth response. [(Khandekar et al., 2024)](https://arxiv.org/abs/2406.12036).\n\nMedCalc Accuracy: Comparison based on category. Exact match for categories risk, severity and diagnosis. Check if within range for the other categories.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "MedCalc Accuracy",
          "run_group": "MedCalc-Bench"
        }
      },
      {
        "value": "MTSamples - Jury Score",
        "description": "MTSamples Replicate is a benchmark that provides transcribed medical reports from various specialties. It is used to evaluate a model's ability to generate clinically appropriate treatment plans based on unstructured patient documentation [(MTSamples, 2025)](https://mtsamples.com).\n\nMTSamples Replicate Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "MTSamples"
        }
      },
      {
        "value": "Medec - MedecFlagAcc",
        "description": "Medec is a benchmark composed of clinical narratives that include either correct documentation or medical errors. Each entry includes sentence-level identifiers and an associated correction task. The model must review the narrative and either identify the erroneous sentence and correct it, or confirm that the text is entirely accurate [(Abacha et al., 2025)](https://arxiv.org/abs/2412.19260).\n\nMedical Error Flag Accuracy: Measures how accurately the model identifies whether a clinical note contains an error (binary classification of correct/incorrect).",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "MedecFlagAcc",
          "run_group": "Medec"
        }
      },
      {
        "value": "HeadQA - EM",
        "description": "HeadQA is a benchmark consisting of biomedical multiple-choice questions intended to evaluate a model's medical knowledge and reasoning. Each instance presents a clinical or scientific question with four answer options, requiring the model to select the most appropriate answer [(Vilares et al., 2019)](https://arxiv.org/abs/1906.04701).\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "HeadQA"
        }
      },
      {
        "value": "Medbullets - EM",
        "description": "Medbullets is a benchmark of USMLE-style medical questions designed to assess a model's ability to understand and apply clinical knowledge. Each question is accompanied by a patient scenario and five multiple-choice options, similar to those found on Step 2 and Step 3 board exams [(MedBullets, 2025)](https://step2.medbullets.com).\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "Medbullets"
        }
      },
      {
        "value": "MedQA - EM",
        "description": "MedQA is an open domain question answering dataset composed of questions from professional medical board exams ([Jin et al. 2020](https://arxiv.org/pdf/2009.13081.pdf)).\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "MedQA"
        }
      },
      {
        "value": "MedMCQA - EM",
        "description": "MedMCQA is a multiple-choice question answering (MCQA) dataset designed to address real-world medical entrance exam questions ([Flores et al. 2020](https://arxiv.org/abs/2203.14371)).\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "MedMCQA"
        }
      },
      {
        "value": "ACI-Bench - Jury Score",
        "description": "ACI-Bench is a benchmark of real-world patient-doctor conversations paired with structured clinical notes. The benchmark evaluates a model's ability to understand spoken medical dialogue and convert it into formal clinical documentation, covering sections such as history of present illness, physical exam findings, results, and assessment and plan [(Yim et al., 2024)](https://www.nature.com/articles/s41597-023-02487-3).\n\nACI-Bench Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "ACI-Bench"
        }
      },
      {
        "value": "MTSamples Procedures - Jury Score",
        "description": "MTSamples Procedures is a benchmark composed of transcribed operative notes, focused on documenting surgical procedures. Each example presents a brief patient case involving a surgical intervention, and the model is tasked with generating a coherent and clinically accurate procedural summary or treatment plan.\n\nMTSamples Procedures Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "MTSamples Procedures"
        }
      },
      {
        "value": "MIMIC-RRS - Jury Score",
        "description": "MIMIC-RRS is a benchmark constructed from radiology reports in the MIMIC-III database. It contains pairs of \u2018Findings\u2018 and \u2018Impression\u2018 sections, enabling evaluation of a model's ability to summarize diagnostic imaging observations into concise, clinically relevant conclusions [(Chen et al., 2023)](https://arxiv.org/abs/2211.08584).\n\nMIMIC-RRS Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "MIMIC-RRS"
        }
      },
      {
        "value": "MIMIC-BHC - Jury Score",
        "description": "MIMIC-BHC is a benchmark focused on summarization of discharge notes into Brief Hospital Course (BHC) sections. It consists of curated discharge notes from MIMIC-IV, each paired with its corresponding BHC summary. The benchmark evaluates a model's ability to condense detailed clinical information into accurate, concise summaries that reflect the patient's hospital stay [(Aali et al., 2024)](https://doi.org/10.1093/jamia/ocae312).\n\nMIMIC-BHC Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "MIMIC-BHC"
        }
      },
      {
        "value": "MedicationQA - Jury Score",
        "description": "MedicationQA is a benchmark composed of open-ended consumer health questions specifically focused on medications. Each example consists of a free-form question and a corresponding medically grounded answer. The benchmark evaluates a model's ability to provide accurate, accessible, and informative medication-related responses for a lay audience.\n\nMedicationQA Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "MedicationQA"
        }
      },
      {
        "value": "PatientInstruct - Jury Score",
        "description": "PatientInstruct is a benchmark designed to evaluate models on generating personalized post-procedure instructions for patients. It includes real-world clinical case details, such as diagnosis, planned procedures, and history and physical notes, from which models must produce clear, actionable instructions appropriate for patients recovering from medical interventions.\n\nPatientInstruct Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "PatientInstruct"
        }
      },
      {
        "value": "MedDialog - Jury Score",
        "description": "MedDialog is a benchmark of real-world doctor-patient conversations focused on health-related concerns and advice. Each dialogue is paired with a one-sentence summary that reflects the core patient question or exchange. The benchmark evaluates a model's ability to condense medical dialogue into concise, informative summaries.\n\nMedDialog Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "MedDialog"
        }
      },
      {
        "value": "MedConfInfo - EM",
        "description": "MedConfInfo is a benchmark comprising clinical notes from adolescent patients. It is used to evaluate whether the content contains sensitive protected health information (PHI) that should be restricted from parental access, in accordance with adolescent confidentiality policies in clinical care. [(Rabbani et al., 2024)](https://jamanetwork.com/journals/jamapediatrics/fullarticle/2814109).\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "MedConfInfo"
        }
      },
      {
        "value": "MEDIQA - Jury Score",
        "description": "MEDIQA is a benchmark designed to evaluate a model's ability to retrieve and generate medically accurate answers to patient-generated questions. Each instance includes a consumer health question, a set of candidate answers (used in ranking tasks), relevance annotations, and optionally, additional context. The benchmark focuses on supporting patient understanding and accessibility in health communication.\n\nMediQA Jury Score: Measures the average score assigned by an LLM-based jury evaluating task performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "Jury Score",
          "run_group": "MEDIQA"
        }
      },
      {
        "value": "ProxySender - EM",
        "description": "ProxySender is a benchmark composed of patient portal messages received by clinicians. It evaluates whether the message was sent by the patient or by a proxy user (e.g., parent, spouse), which is critical for understanding who is communicating with healthcare providers. [(Tse G, et al., 2025)](https://doi.org/10.1001/jamapediatrics.2024.4438).\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "ProxySender"
        }
      },
      {
        "value": "PubMedQA - EM",
        "description": "PubMedQA is a biomedical question-answering dataset that evaluates a model's ability to interpret scientific literature. It consists of PubMed abstracts paired with yes/no/maybe questions derived from the content. The benchmark assesses a model's capability to reason over biomedical texts and provide factually grounded answers.\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "PubMedQA"
        }
      },
      {
        "value": "EHRSQL - EHRSQLExeAcc",
        "description": "EHRSQL is a benchmark designed to evaluate models on generating structured queries for clinical research. Each example includes a natural language question and a database schema, and the task is to produce an SQL query that would return the correct result for a biomedical research objective. This benchmark assesses a model's understanding of medical terminology, data structures, and query construction.\n\nExecution accuracy for Generated Query: Measures the proportion of correctly predicted answerable questions among all questions predicted to be answerable.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EHRSQLExeAcc",
          "run_group": "EHRSQL"
        }
      },
      {
        "value": "BMT-Status - EM",
        "description": "BMT-Status is a benchmark composed of clinical notes and associated binary questions related to bone marrow transplant (BMT), hematopoietic stem cell transplant (HSCT), or hematopoietic cell transplant (HCT) status. The goal is to determine whether the patient received a subsequent transplant based on the provided clinical documentation.\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "BMT-Status"
        }
      },
      {
        "value": "RaceBias - EM",
        "description": "RaceBias is a benchmark used to evaluate language models for racially biased or inappropriate content in medical question-answering scenarios. Each instance consists of a medical question and a model-generated response. The task is to classify whether the response contains race-based, harmful, or inaccurate content. This benchmark supports research into bias detection and fairness in clinical AI systems.\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "RaceBias"
        }
      },
      {
        "value": "N2C2-CT - EM",
        "description": "A dataset that provides clinical notes and asks the model to classify whether the patient is a valid candidate for a provided clinical trial.\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "N2C2-CT"
        }
      },
      {
        "value": "MedHallu - EM",
        "description": "MedHallu is a benchmark focused on evaluating factual correctness in biomedical question answering. Each instance contains a PubMed-derived knowledge snippet, a biomedical question, and a model-generated answer. The task is to classify whether the answer is factually correct or contains hallucinated (non-grounded) information. This benchmark is designed to assess the factual reliability of medical language models.\n\nExact match: Fraction of instances that the predicted output matches a correct reference exactly.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "EM",
          "run_group": "MedHallu"
        }
      },
      {
        "value": "MIMIC-IV Billing Code - MIMICBillingF1",
        "description": "MIMIC-IV Billing Code is a benchmark derived from discharge summaries in the MIMIC-IV database, paired with their corresponding ICD-10 billing codes. The task requires models to extract structured billing codes based on free-text clinical notes, reflecting real-world hospital coding tasks for financial reimbursement.\n\nF1 Score for MIMIC Billing Codes: Measures the harmonic mean of precision and recall for ICD codes, providing a balanced evaluation of the model's performance.",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {
          "metric": "MIMICBillingF1",
          "run_group": "MIMIC-IV Billing Code"
        }
      }
    ],
    "rows": [
      [
        {
          "value": "Claude 4.6 Opus",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.45625,
          "markdown": false
        },
        {
          "value": 0.22272727272727272,
          "description": "min=0.223, mean=0.223, max=0.223, sum=0.223 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4.371584699453542,
          "description": "min=4.372, mean=4.372, max=4.372, sum=4.372 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.5695142378559463,
          "description": "min=0.57, mean=0.57, max=0.57, sum=0.57 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.83956121867538,
          "description": "min=0.553, mean=0.84, max=0.921, sum=4.198 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=chemistry,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=medicine,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=pharmacology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=psychology,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.4577922077922078,
          "description": "min=0.458, mean=0.458, max=0.458, sum=0.458 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.9300864100549883,
          "description": "min=0.93, mean=0.93, max=0.93, sum=0.93 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.8137700215156586,
          "description": "min=0.814, mean=0.814, max=0.814, sum=0.814 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4.911111111111111,
          "description": "min=4.911, mean=4.911, max=4.911, sum=4.911 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4.088541666666664,
          "description": "min=4.089, mean=4.089, max=4.089, sum=4.089 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4.122270742358076,
          "description": "min=4.122, mean=4.122, max=4.122, sum=4.122 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 3.763333333333333,
          "description": "min=3.763, mean=3.763, max=3.763, sum=3.763 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4.49927431059506,
          "description": "min=4.499, mean=4.499, max=4.499, sum=4.499 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4.994933333333333,
          "description": "min=4.995, mean=4.995, max=4.995, sum=4.995 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 3.868335349635368,
          "description": "min=3.783, mean=3.868, max=3.954, sum=7.737 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-opus-4.6",
            "med_dialog,subset=icliniq:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.6788,
          "description": "min=0.679, mean=0.679, max=0.679, sum=0.679 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4.793333333333332,
          "description": "min=4.793, mean=4.793, max=4.793, sum=4.793 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.99,
          "description": "min=0.99, mean=0.99, max=0.99, sum=0.99 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.733,
          "description": "min=0.733, mean=0.733, max=0.733, sum=0.733 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.26469577174286696,
          "description": "min=0.265, mean=0.265, max=0.265, sum=0.265 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.864,
          "description": "min=0.864, mean=0.864, max=0.864, sum=0.864 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.7664670658682635,
          "description": "min=0.766, mean=0.766, max=0.766, sum=0.766 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.7471094190637336,
          "description": "min=0.686, mean=0.747, max=0.821, sum=2.241 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.9395,
          "description": "min=0.94, mean=0.94, max=0.94, sum=0.94 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.3561229455745828,
          "description": "min=0.356, mean=0.356, max=0.356, sum=0.356 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=anthropic_claude-opus-4.6"
          ]
        }
      ],
      [
        {
          "value": "Gemini 3.1 Pro (Preview)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.6520833333333333,
          "style": {
            "font-weight": "bold"
          },
          "markdown": false
        },
        {
          "value": 0.49454545454545457,
          "description": "min=0.495, mean=0.495, max=0.495, sum=0.495 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4.123341139734576,
          "description": "min=4.123, mean=4.123, max=4.123, sum=4.123 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.6800670016750419,
          "description": "min=0.68, mean=0.68, max=0.68, sum=0.68 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.9202263426831209,
          "description": "min=0.669, mean=0.92, max=0.991, sum=4.601 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=medicine,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=psychology,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.7662337662337663,
          "description": "min=0.766, mean=0.766, max=0.766, sum=0.766 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.9638648860958366,
          "description": "min=0.964, mean=0.964, max=0.964, sum=0.964 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.8737748027731294,
          "description": "min=0.874, mean=0.874, max=0.874, sum=0.874 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4.908333333333332,
          "description": "min=4.908, mean=4.908, max=4.908, sum=4.908 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4.177083333333329,
          "description": "min=4.177, mean=4.177, max=4.177, sum=4.177 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 3.8131228305900775,
          "description": "min=3.813, mean=3.813, max=3.813, sum=3.813 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 3.709999999999999,
          "description": "min=3.71, mean=3.71, max=3.71, sum=3.71 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4.5258829221093295,
          "description": "min=4.526, mean=4.526, max=4.526, sum=4.526 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4.885933333333378,
          "description": "min=4.886, mean=4.886, max=4.886, sum=4.886 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 3.5520290433290524,
          "description": "min=3.42, mean=3.552, max=3.684, sum=7.104 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.1-pro-preview",
            "med_dialog,subset=icliniq:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.6996,
          "description": "min=0.7, mean=0.7, max=0.7, sum=0.7 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4.844444444444442,
          "description": "min=4.844, mean=4.844, max=4.844, sum=4.844 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.9888,
          "description": "min=0.989, mean=0.989, max=0.989, sum=0.989 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.817,
          "description": "min=0.817, mean=0.817, max=0.817, sum=0.817 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.44345135785493295,
          "description": "min=0.443, mean=0.443, max=0.443, sum=0.443 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.887,
          "description": "min=0.887, mean=0.887, max=0.887, sum=0.887 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.9161676646706587,
          "description": "min=0.916, mean=0.916, max=0.916, sum=0.916 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.765721940214326,
          "description": "min=0.66, mean=0.766, max=0.9, sum=2.297 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.948,
          "description": "min=0.948, mean=0.948, max=0.948, sum=0.948 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.4726577473121329,
          "description": "min=0.473, mean=0.473, max=0.473, sum=0.473 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=google_gemini-3.1-pro-preview"
          ]
        }
      ],
      [
        {
          "value": "Gemini-3.5-flash",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.6416666666666667,
          "markdown": false
        },
        {
          "value": 0.5045454545454545,
          "description": "min=0.505, mean=0.505, max=0.505, sum=0.505 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4.217798594847767,
          "description": "min=4.218, mean=4.218, max=4.218, sum=4.218 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.6432160804020101,
          "description": "min=0.643, mean=0.643, max=0.643, sum=0.643 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.9226398521356624,
          "description": "min=0.677, mean=0.923, max=0.987, sum=4.613 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=medicine,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=psychology,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.7987012987012987,
          "description": "min=0.799, mean=0.799, max=0.799, sum=0.799 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.9591516103692066,
          "description": "min=0.959, mean=0.959, max=0.959, sum=0.959 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.8596700932345207,
          "description": "min=0.86, mean=0.86, max=0.86, sum=0.86 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4.886111111111108,
          "description": "min=4.886, mean=4.886, max=4.886, sum=4.886 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4.161458333333328,
          "description": "min=4.161, mean=4.161, max=4.161, sum=4.161 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.872018810883436,
          "description": "min=3.872, mean=3.872, max=3.872, sum=3.872 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.73,
          "description": "min=3.73, mean=3.73, max=3.73, sum=3.73 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4.571843251088528,
          "description": "min=4.572, mean=4.572, max=4.572, sum=4.572 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4.9222000000000214,
          "description": "min=4.922, mean=4.922, max=4.922, sum=4.922 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.570790690690698,
          "description": "min=3.465, mean=3.571, max=3.677, sum=7.142 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.5-flash",
            "med_dialog,subset=icliniq:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.699,
          "description": "min=0.699, mean=0.699, max=0.699, sum=0.699 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4.855555555555554,
          "description": "min=4.856, mean=4.856, max=4.856, sum=4.856 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.9702,
          "description": "min=0.97, mean=0.97, max=0.97, sum=0.97 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.802,
          "description": "min=0.802, mean=0.802, max=0.802, sum=0.802 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.21725678927466482,
          "description": "min=0.217, mean=0.217, max=0.217, sum=0.217 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.8916,
          "description": "min=0.892, mean=0.892, max=0.892, sum=0.892 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.9221556886227545,
          "description": "min=0.922, mean=0.922, max=0.922, sum=0.922 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.7600817822899041,
          "description": "min=0.673, mean=0.76, max=0.873, sum=2.28 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.9185,
          "description": "min=0.918, mean=0.918, max=0.918, sum=0.918 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.46945897835887834,
          "description": "min=0.469, mean=0.469, max=0.469, sum=0.469 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=google_gemini-3.5-flash"
          ]
        }
      ],
      [
        {
          "value": "GPT-5.4 (2026-03-05)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.5375,
          "markdown": false
        },
        {
          "value": 0.4672727272727273,
          "description": "min=0.467, mean=0.467, max=0.467, sum=0.467 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.463700234192029,
          "description": "min=4.464, mean=4.464, max=4.464, sum=4.464 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.5946398659966499,
          "description": "min=0.595, mean=0.595, max=0.595, sum=0.595 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.9154686257662366,
          "description": "min=0.687, mean=0.915, max=0.985, sum=4.577 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.8733766233766234,
          "description": "min=0.873, mean=0.873, max=0.873, sum=0.873 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.9528672427336999,
          "description": "min=0.953, mean=0.953, max=0.953, sum=0.953 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.8166387759980875,
          "description": "min=0.817, mean=0.817, max=0.817, sum=0.817 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.913888888888889,
          "description": "min=4.914, mean=4.914, max=4.914, sum=4.914 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.2760416666666625,
          "description": "min=4.276, mean=4.276, max=4.276, sum=4.276 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.021946030679651,
          "description": "min=4.022, mean=4.022, max=4.022, sum=4.022 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 3.6933333333333325,
          "description": "min=3.693, mean=3.693, max=3.693, sum=3.693 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.396226415094335,
          "description": "min=4.396, mean=4.396, max=4.396, sum=4.396 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.930533333333357,
          "description": "min=4.931, mean=4.931, max=4.931, sum=4.931 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 3.840635049335062,
          "description": "min=3.765, mean=3.841, max=3.916, sum=7.681 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.6968,
          "description": "min=0.697, mean=0.697, max=0.697, sum=0.697 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.755555555555554,
          "description": "min=4.756, mean=4.756, max=4.756, sum=4.756 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.9728,
          "description": "min=0.973, mean=0.973, max=0.973, sum=0.973 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.746,
          "description": "min=0.746, mean=0.746, max=0.746, sum=0.746 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.3437607425232039,
          "description": "min=0.344, mean=0.344, max=0.344, sum=0.344 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.8998,
          "description": "min=0.9, mean=0.9, max=0.9, sum=0.9 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.6526946107784432,
          "description": "min=0.653, mean=0.653, max=0.653, sum=0.653 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.7547941342357586,
          "description": "min=0.638, mean=0.755, max=0.905, sum=2.264 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.7645,
          "description": "min=0.764, mean=0.764, max=0.764, sum=0.764 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.4882719524691857,
          "description": "min=0.488, mean=0.488, max=0.488, sum=0.488 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=openai_gpt-5.4"
          ]
        }
      ],
      [
        {
          "value": "GPT-5.4 mini",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.5520833333333334,
          "markdown": false
        },
        {
          "value": 0.4781818181818182,
          "description": "min=0.478, mean=0.478, max=0.478, sum=0.478 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.319281811085087,
          "description": "min=4.319, mean=4.319, max=4.319, sum=4.319 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.7219430485762144,
          "description": "min=0.722, mean=0.722, max=0.722, sum=0.722 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.9141952545720159,
          "description": "min=0.689, mean=0.914, max=0.98, sum=4.571 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.8733766233766234,
          "description": "min=0.873, mean=0.873, max=0.873, sum=0.873 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.9528672427336999,
          "description": "min=0.953, mean=0.953, max=0.953, sum=0.953 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.8230934735835524,
          "description": "min=0.823, mean=0.823, max=0.823, sum=0.823 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.897222222222221,
          "description": "min=4.897, mean=4.897, max=4.897, sum=4.897 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.174479166666661,
          "description": "min=4.174, mean=4.174, max=4.174, sum=4.174 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.155189788377547,
          "description": "min=4.155, mean=4.155, max=4.155, sum=4.155 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 3.6133333333333324,
          "description": "min=3.613, mean=3.613, max=3.613, sum=3.613 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.3280116110304725,
          "description": "min=4.328, mean=4.328, max=4.328, sum=4.328 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.905133333333364,
          "description": "min=4.905, mean=4.905, max=4.905, sum=4.905 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 3.831052252252272,
          "description": "min=3.755, mean=3.831, max=3.908, sum=7.662 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4-mini",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.6982,
          "description": "min=0.698, mean=0.698, max=0.698, sum=0.698 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.697777777777776,
          "description": "min=4.698, mean=4.698, max=4.698, sum=4.698 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.8896,
          "description": "min=0.89, mean=0.89, max=0.89, sum=0.89 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.764,
          "description": "min=0.764, mean=0.764, max=0.764, sum=0.764 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.34444826400825024,
          "description": "min=0.344, mean=0.344, max=0.344, sum=0.344 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.8928,
          "description": "min=0.893, mean=0.893, max=0.893, sum=0.893 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.7125748502994012,
          "description": "min=0.713, mean=0.713, max=0.713, sum=0.713 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.7552171460800903,
          "description": "min=0.643, mean=0.755, max=0.889, sum=2.266 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.8985,
          "description": "min=0.898, mean=0.898, max=0.898, sum=0.898 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.4730430206016039,
          "description": "min=0.473, mean=0.473, max=0.473, sum=0.473 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=openai_gpt-5.4-mini"
          ]
        }
      ],
      [
        {
          "value": "Claude 3.7 Sonnet (20250219)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.45,
          "markdown": false
        },
        {
          "value": 0.21,
          "description": "min=0.21, mean=0.21, max=0.21, sum=0.21 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 4.425969294821761,
          "description": "min=4.426, mean=4.426, max=4.426, sum=4.426 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.628140703517588,
          "description": "min=0.628, mean=0.628, max=0.628, sum=0.628 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.912,
          "description": "min=0.912, mean=0.912, max=0.912, sum=0.912 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.6493506493506493,
          "description": "min=0.649, mean=0.649, max=0.649, sum=0.649 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.856858846918489,
          "description": "min=0.857, mean=0.857, max=0.857, sum=0.857 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.74,
          "description": "min=0.74, mean=0.74, max=0.74, sum=0.74 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 4.617592592592595,
          "description": "min=4.618, mean=4.618, max=4.618, sum=4.618 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 4.0234375,
          "description": "min=4.023, mean=4.023, max=4.023, sum=4.023 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 4.5008888888888885,
          "description": "min=4.501, mean=4.501, max=4.501, sum=4.501 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 3.9911111111111097,
          "description": "min=3.991, mean=3.991, max=3.991, sum=3.991 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 4.628124496049018,
          "description": "min=4.628, mean=4.628, max=4.628, sum=4.628 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 3.6611265004616884,
          "description": "min=3.661, mean=3.661, max=3.661, sum=3.661 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 4.197555555555553,
          "description": "min=4.146, mean=4.198, max=4.249, sum=8.395 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "med_dialog,subset=icliniq:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.811,
          "description": "min=0.811, mean=0.811, max=0.811, sum=0.811 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 4.72814814814815,
          "description": "min=4.728, mean=4.728, max=4.728, sum=4.728 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.62,
          "description": "min=0.62, mean=0.62, max=0.62, sum=0.62 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.625,
          "description": "min=0.625, mean=0.625, max=0.625, sum=0.625 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.081,
          "description": "min=0.081, mean=0.081, max=0.081, sum=0.081 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.9,
          "description": "min=0.9, mean=0.9, max=0.9, sum=0.9 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.6586826347305389,
          "description": "min=0.659, mean=0.659, max=0.659, sum=0.659 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.813953488372093,
          "description": "min=0.756, mean=0.814, max=0.895, sum=2.442 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.877,
          "description": "min=0.877, mean=0.877, max=0.877, sum=0.877 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.3554857455069554,
          "description": "min=0.355, mean=0.355, max=0.355, sum=0.355 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        }
      ],
      [
        {
          "value": "DeepSeek R1",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.48541666666666666,
          "markdown": false
        },
        {
          "value": 0.348,
          "description": "min=0.348, mean=0.348, max=0.348, sum=0.348 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.533177205308355,
          "description": "min=4.533, mean=4.533, max=4.533, sum=4.533 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.5912897822445561,
          "description": "min=0.591, mean=0.591, max=0.591, sum=0.591 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.721,
          "description": "min=0.721, mean=0.721, max=0.721, sum=0.721 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.6558441558441559,
          "description": "min=0.656, mean=0.656, max=0.656, sum=0.656 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.856858846918489,
          "description": "min=0.857, mean=0.857, max=0.857, sum=0.857 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.745,
          "description": "min=0.745, mean=0.745, max=0.745, sum=0.745 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.673148148148149,
          "description": "min=4.673, mean=4.673, max=4.673, sum=4.673 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.230902777777779,
          "description": "min=4.231, mean=4.231, max=4.231, sum=4.231 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.610499999999993,
          "description": "min=4.61, mean=4.61, max=4.61, sum=4.61 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 3.9188888888888855,
          "description": "min=3.919, mean=3.919, max=3.919, sum=3.919 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.584744396065148,
          "description": "min=4.585, mean=4.585, max=4.585, sum=4.585 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.784856879039718,
          "description": "min=4.785, mean=4.785, max=4.785, sum=4.785 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.435555555555553,
          "description": "min=4.361, mean=4.436, max=4.51, sum=8.871 (2)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.739,
          "description": "min=0.739, mean=0.739, max=0.739, sum=0.739 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.757777777777781,
          "description": "min=4.758, mean=4.758, max=4.758, sum=4.758 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.7433333333333333,
          "description": "min=0.743, mean=0.743, max=0.743, sum=0.743 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.743,
          "description": "min=0.743, mean=0.743, max=0.743, sum=0.743 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.076,
          "description": "min=0.076, mean=0.076, max=0.076, sum=0.076 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.7772727272727272,
          "description": "min=0.777, mean=0.777, max=0.777, sum=0.777 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.9161676646706587,
          "description": "min=0.916, mean=0.916, max=0.916, sum=0.916 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.8643410852713179,
          "description": "min=0.779, mean=0.864, max=0.93, sum=2.593 (3)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.847,
          "description": "min=0.847, mean=0.847, max=0.847, sum=0.847 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.2989618858066198,
          "description": "min=0.299, mean=0.299, max=0.299, sum=0.299 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        }
      ],
      [
        {
          "value": "Gemini 2.0 Flash",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.3416666666666667,
          "markdown": false
        },
        {
          "value": 0.158,
          "description": "min=0.158, mean=0.158, max=0.158, sum=0.158 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 4.116835805360402,
          "description": "min=4.117, mean=4.117, max=4.117, sum=4.117 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.5963149078726968,
          "description": "min=0.596, mean=0.596, max=0.596, sum=0.596 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.88,
          "description": "min=0.88, mean=0.88, max=0.88, sum=0.88 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.6298701298701299,
          "description": "min=0.63, mean=0.63, max=0.63, sum=0.63 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.8489065606361829,
          "description": "min=0.849, mean=0.849, max=0.849, sum=0.849 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.748,
          "description": "min=0.748, mean=0.748, max=0.748, sum=0.748 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 4.612037037037038,
          "description": "min=4.612, mean=4.612, max=4.612, sum=4.612 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 3.641493055555557,
          "description": "min=3.641, mean=3.641, max=3.641, sum=3.641 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 4.248500000000007,
          "description": "min=4.249, mean=4.249, max=4.249, sum=4.249 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 3.8299999999999987,
          "description": "min=3.83, mean=3.83, max=3.83, sum=3.83 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 4.428801806160294,
          "description": "min=4.429, mean=4.429, max=4.429, sum=4.429 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 3.4361341951369657,
          "description": "min=3.436, mean=3.436, max=3.436, sum=3.436 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 4.224527777777782,
          "description": "min=4.178, mean=4.225, max=4.271, sum=8.449 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "med_dialog,subset=icliniq:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.75,
          "description": "min=0.75, mean=0.75, max=0.75, sum=0.75 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 4.767407407407409,
          "description": "min=4.767, mean=4.767, max=4.767, sum=4.767 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.7466666666666667,
          "description": "min=0.747, mean=0.747, max=0.747, sum=0.747 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.64,
          "description": "min=0.64, mean=0.64, max=0.64, sum=0.64 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.165,
          "description": "min=0.165, mean=0.165, max=0.165, sum=0.165 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.7954545454545454,
          "description": "min=0.795, mean=0.795, max=0.795, sum=0.795 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.8562874251497006,
          "description": "min=0.856, mean=0.856, max=0.856, sum=0.856 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.7170542635658914,
          "description": "min=0.523, mean=0.717, max=0.93, sum=2.151 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.908,
          "description": "min=0.908, mean=0.908, max=0.908, sum=0.908 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.23228162127907237,
          "description": "min=0.232, mean=0.232, max=0.232, sum=0.232 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        }
      ],
      [
        {
          "value": "Gemini 2.5 Pro (05-06 preview)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.5291666666666667,
          "markdown": false
        },
        {
          "value": 0.348,
          "description": "min=0.348, mean=0.348, max=0.348, sum=0.348 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 4.66575591985429,
          "description": "min=4.666, mean=4.666, max=4.666, sum=4.666 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.5477386934673367,
          "description": "min=0.548, mean=0.548, max=0.548, sum=0.548 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.931,
          "description": "min=0.931, mean=0.931, max=0.931, sum=0.931 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.8084415584415584,
          "description": "min=0.808, mean=0.808, max=0.808, sum=0.808 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.9343936381709742,
          "description": "min=0.934, mean=0.934, max=0.934, sum=0.934 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.834,
          "description": "min=0.834, mean=0.834, max=0.834, sum=0.834 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 4.732407407407409,
          "description": "min=4.732, mean=4.732, max=4.732, sum=4.732 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 4.276909722222224,
          "description": "min=4.277, mean=4.277, max=4.277, sum=4.277 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 4.25116666666667,
          "description": "min=4.251, mean=4.251, max=4.251, sum=4.251 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 3.756666666666665,
          "description": "min=3.757, mean=3.757, max=3.757, sum=3.757 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 4.837768101919036,
          "description": "min=4.838, mean=4.838, max=4.838, sum=4.838 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 4.8710372422283905,
          "description": "min=4.871, mean=4.871, max=4.871, sum=4.871 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 3.7637777777777677,
          "description": "min=3.636, mean=3.764, max=3.892, sum=7.528 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.677,
          "description": "min=0.677, mean=0.677, max=0.677, sum=0.677 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 4.966666666666669,
          "description": "min=4.967, mean=4.967, max=4.967, sum=4.967 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.7366666666666667,
          "description": "min=0.737, mean=0.737, max=0.737, sum=0.737 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.764,
          "description": "min=0.764, mean=0.764, max=0.764, sum=0.764 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.213,
          "description": "min=0.213, mean=0.213, max=0.213, sum=0.213 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.7727272727272727,
          "description": "min=0.773, mean=0.773, max=0.773, sum=0.773 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.8263473053892215,
          "description": "min=0.826, mean=0.826, max=0.826, sum=0.826 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.8062015503875969,
          "description": "min=0.698, mean=0.806, max=0.895, sum=2.419 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.864,
          "description": "min=0.864, mean=0.864, max=0.864, sum=0.864 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.2620596169954098,
          "description": "min=0.262, mean=0.262, max=0.262, sum=0.262 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        }
      ],
      [
        {
          "value": "Llama 3.3 Instruct (70B)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.23333333333333334,
          "markdown": false
        },
        {
          "value": 0.113,
          "description": "min=0.113, mean=0.113, max=0.113, sum=0.113 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 4.144418423106954,
          "description": "min=4.144, mean=4.144, max=4.144, sum=4.144 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.5293132328308208,
          "description": "min=0.529, mean=0.529, max=0.529, sum=0.529 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.854,
          "description": "min=0.854, mean=0.854, max=0.854, sum=0.854 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.6071428571428571,
          "description": "min=0.607, mean=0.607, max=0.607, sum=0.607 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.8011928429423459,
          "description": "min=0.801, mean=0.801, max=0.801, sum=0.801 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.724,
          "description": "min=0.724, mean=0.724, max=0.724, sum=0.724 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 4.124074074074071,
          "description": "min=4.124, mean=4.124, max=4.124, sum=4.124 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 3.8376736111111147,
          "description": "min=3.838, mean=3.838, max=3.838, sum=3.838 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 4.058111111111116,
          "description": "min=4.058, mean=4.058, max=4.058, sum=4.058 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 3.7677777777777766,
          "description": "min=3.768, mean=3.768, max=3.768, sum=3.768 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 4.490566037735847,
          "description": "min=4.491, mean=4.491, max=4.491, sum=4.491 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 3.5866420437057647,
          "description": "min=3.587, mean=3.587, max=3.587, sum=3.587 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 4.086388888888882,
          "description": "min=4.039, mean=4.086, max=4.133, sum=8.173 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "med_dialog,subset=icliniq:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.789,
          "description": "min=0.789, mean=0.789, max=0.789, sum=0.789 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 4.674074074074076,
          "description": "min=4.674, mean=4.674, max=4.674, sum=4.674 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.6833333333333333,
          "description": "min=0.683, mean=0.683, max=0.683, sum=0.683 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.733,
          "description": "min=0.733, mean=0.733, max=0.733, sum=0.733 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.074,
          "description": "min=0.074, mean=0.074, max=0.074, sum=0.074 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.8363636363636363,
          "description": "min=0.836, mean=0.836, max=0.836, sum=0.836 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.5688622754491018,
          "description": "min=0.569, mean=0.569, max=0.569, sum=0.569 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.813953488372093,
          "description": "min=0.744, mean=0.814, max=0.884, sum=2.442 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=CREATININE,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.882,
          "description": "min=0.882, mean=0.882, max=0.882, sum=0.882 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.19714220046154962,
          "description": "min=0.197, mean=0.197, max=0.197, sum=0.197 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        }
      ],
      [
        {
          "value": "Muse Spark (2026-04-08)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.6208333333333333,
          "markdown": false
        },
        {
          "value": 0.48,
          "description": "min=0.48, mean=0.48, max=0.48, sum=0.48 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4.60499609679937,
          "description": "min=4.605, mean=4.605, max=4.605, sum=4.605 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.6398659966499163,
          "description": "min=0.64, mean=0.64, max=0.64, sum=0.64 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.922817085384881,
          "description": "min=0.684, mean=0.923, max=0.993, sum=4.614 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=chemistry,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=medicine,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=pharmacology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=psychology,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.8571428571428571,
          "description": "min=0.857, mean=0.857, max=0.857, sum=0.857 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.9512961508248232,
          "description": "min=0.951, mean=0.951, max=0.951, sum=0.951 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.856562275878556,
          "description": "min=0.857, mean=0.857, max=0.857, sum=0.857 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4.969444444444444,
          "description": "min=4.969, mean=4.969, max=4.969, sum=4.969 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4.359374999999998,
          "description": "min=4.359, mean=4.359, max=4.359, sum=4.359 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4.153846153846153,
          "description": "min=4.154, mean=4.154, max=4.154, sum=4.154 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4.0266666666666655,
          "description": "min=4.027, mean=4.027, max=4.027, sum=4.027 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4.613449443638116,
          "description": "min=4.613, mean=4.613, max=4.613, sum=4.613 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4.903733333333328,
          "description": "min=4.904, mean=4.904, max=4.904, sum=4.904 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 3.818098755898764,
          "description": "min=3.762, mean=3.818, max=3.874, sum=7.636 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_quiet_sand_v1alpha",
            "med_dialog,subset=icliniq:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.6948,
          "description": "min=0.695, mean=0.695, max=0.695, sum=0.695 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4.8488888888888875,
          "description": "min=4.849, mean=4.849, max=4.849, sum=4.849 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.5976,
          "description": "min=0.598, mean=0.598, max=0.598, sum=0.598 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.776,
          "description": "min=0.776, mean=0.776, max=0.776, sum=0.776 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.12048814025438294,
          "description": "min=0.049, mean=0.12, max=0.192, sum=0.241 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.68,
          "description": "min=0.68, mean=0.68, max=0.68, sum=0.68 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.8323353293413174,
          "description": "min=0.832, mean=0.832, max=0.832, sum=0.832 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.7597292724196277,
          "description": "min=0.68, mean=0.76, max=0.864, sum=2.279 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=CREATININE,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.8955,
          "description": "min=0.895, mean=0.895, max=0.895, sum=0.895 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.36521697665286573,
          "description": "min=0.365, mean=0.365, max=0.365, sum=0.365 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=meta_quiet_sand_v1alpha"
          ]
        }
      ]
    ],
    "links": [
      {
        "text": "LaTeX",
        "href": "benchmarks/releases/v5.0.0/groups/latex/medhelm_scenarios_accuracy.tex"
      },
      {
        "text": "JSON",
        "href": "benchmarks/releases/v5.0.0/groups/json/medhelm_scenarios_accuracy.json"
      }
    ],
    "name": "accuracy"
  },
  {
    "title": "Efficiency",
    "header": [
      {
        "value": "Model",
        "markdown": false,
        "metadata": {}
      },
      {
        "value": "Mean win rate",
        "description": "How many models this model outperforms on average (over columns).",
        "markdown": false,
        "lower_is_better": false,
        "metadata": {}
      },
      {
        "value": "MedCalc-Bench - Observed inference time (s)",
        "description": "MedCalc-Bench is a benchmark designed to evaluate models on their ability to compute clinically relevant values from patient notes. Each instance consists of a clinical note describing the patient's condition, a diagnostic question targeting a specific medical value, and a ground truth response. [(Khandekar et al., 2024)](https://arxiv.org/abs/2406.12036).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MedCalc-Bench"
        }
      },
      {
        "value": "MTSamples - Observed inference time (s)",
        "description": "MTSamples Replicate is a benchmark that provides transcribed medical reports from various specialties. It is used to evaluate a model's ability to generate clinically appropriate treatment plans based on unstructured patient documentation [(MTSamples, 2025)](https://mtsamples.com).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MTSamples"
        }
      },
      {
        "value": "Medec - Observed inference time (s)",
        "description": "Medec is a benchmark composed of clinical narratives that include either correct documentation or medical errors. Each entry includes sentence-level identifiers and an associated correction task. The model must review the narrative and either identify the erroneous sentence and correct it, or confirm that the text is entirely accurate [(Abacha et al., 2025)](https://arxiv.org/abs/2412.19260).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "Medec"
        }
      },
      {
        "value": "HeadQA - Observed inference time (s)",
        "description": "HeadQA is a benchmark consisting of biomedical multiple-choice questions intended to evaluate a model's medical knowledge and reasoning. Each instance presents a clinical or scientific question with four answer options, requiring the model to select the most appropriate answer [(Vilares et al., 2019)](https://arxiv.org/abs/1906.04701).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "HeadQA"
        }
      },
      {
        "value": "Medbullets - Observed inference time (s)",
        "description": "Medbullets is a benchmark of USMLE-style medical questions designed to assess a model's ability to understand and apply clinical knowledge. Each question is accompanied by a patient scenario and five multiple-choice options, similar to those found on Step 2 and Step 3 board exams [(MedBullets, 2025)](https://step2.medbullets.com).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "Medbullets"
        }
      },
      {
        "value": "MedQA - Observed inference time (s)",
        "description": "MedQA is an open domain question answering dataset composed of questions from professional medical board exams ([Jin et al. 2020](https://arxiv.org/pdf/2009.13081.pdf)).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MedQA"
        }
      },
      {
        "value": "MedMCQA - Observed inference time (s)",
        "description": "MedMCQA is a multiple-choice question answering (MCQA) dataset designed to address real-world medical entrance exam questions ([Flores et al. 2020](https://arxiv.org/abs/2203.14371)).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MedMCQA"
        }
      },
      {
        "value": "ACI-Bench - Observed inference time (s)",
        "description": "ACI-Bench is a benchmark of real-world patient-doctor conversations paired with structured clinical notes. The benchmark evaluates a model's ability to understand spoken medical dialogue and convert it into formal clinical documentation, covering sections such as history of present illness, physical exam findings, results, and assessment and plan [(Yim et al., 2024)](https://www.nature.com/articles/s41597-023-02487-3).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "ACI-Bench"
        }
      },
      {
        "value": "MTSamples Procedures - Observed inference time (s)",
        "description": "MTSamples Procedures is a benchmark composed of transcribed operative notes, focused on documenting surgical procedures. Each example presents a brief patient case involving a surgical intervention, and the model is tasked with generating a coherent and clinically accurate procedural summary or treatment plan.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MTSamples Procedures"
        }
      },
      {
        "value": "MIMIC-RRS - Observed inference time (s)",
        "description": "MIMIC-RRS is a benchmark constructed from radiology reports in the MIMIC-III database. It contains pairs of \u2018Findings\u2018 and \u2018Impression\u2018 sections, enabling evaluation of a model's ability to summarize diagnostic imaging observations into concise, clinically relevant conclusions [(Chen et al., 2023)](https://arxiv.org/abs/2211.08584).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MIMIC-RRS"
        }
      },
      {
        "value": "MIMIC-BHC - Observed inference time (s)",
        "description": "MIMIC-BHC is a benchmark focused on summarization of discharge notes into Brief Hospital Course (BHC) sections. It consists of curated discharge notes from MIMIC-IV, each paired with its corresponding BHC summary. The benchmark evaluates a model's ability to condense detailed clinical information into accurate, concise summaries that reflect the patient's hospital stay [(Aali et al., 2024)](https://doi.org/10.1093/jamia/ocae312).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MIMIC-BHC"
        }
      },
      {
        "value": "MedicationQA - Observed inference time (s)",
        "description": "MedicationQA is a benchmark composed of open-ended consumer health questions specifically focused on medications. Each example consists of a free-form question and a corresponding medically grounded answer. The benchmark evaluates a model's ability to provide accurate, accessible, and informative medication-related responses for a lay audience.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MedicationQA"
        }
      },
      {
        "value": "PatientInstruct - Observed inference time (s)",
        "description": "PatientInstruct is a benchmark designed to evaluate models on generating personalized post-procedure instructions for patients. It includes real-world clinical case details, such as diagnosis, planned procedures, and history and physical notes, from which models must produce clear, actionable instructions appropriate for patients recovering from medical interventions.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "PatientInstruct"
        }
      },
      {
        "value": "MedDialog - Observed inference time (s)",
        "description": "MedDialog is a benchmark of real-world doctor-patient conversations focused on health-related concerns and advice. Each dialogue is paired with a one-sentence summary that reflects the core patient question or exchange. The benchmark evaluates a model's ability to condense medical dialogue into concise, informative summaries.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MedDialog"
        }
      },
      {
        "value": "MedConfInfo - Observed inference time (s)",
        "description": "MedConfInfo is a benchmark comprising clinical notes from adolescent patients. It is used to evaluate whether the content contains sensitive protected health information (PHI) that should be restricted from parental access, in accordance with adolescent confidentiality policies in clinical care. [(Rabbani et al., 2024)](https://jamanetwork.com/journals/jamapediatrics/fullarticle/2814109).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MedConfInfo"
        }
      },
      {
        "value": "MEDIQA - Observed inference time (s)",
        "description": "MEDIQA is a benchmark designed to evaluate a model's ability to retrieve and generate medically accurate answers to patient-generated questions. Each instance includes a consumer health question, a set of candidate answers (used in ranking tasks), relevance annotations, and optionally, additional context. The benchmark focuses on supporting patient understanding and accessibility in health communication.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MEDIQA"
        }
      },
      {
        "value": "ProxySender - Observed inference time (s)",
        "description": "ProxySender is a benchmark composed of patient portal messages received by clinicians. It evaluates whether the message was sent by the patient or by a proxy user (e.g., parent, spouse), which is critical for understanding who is communicating with healthcare providers. [(Tse G, et al., 2025)](https://doi.org/10.1001/jamapediatrics.2024.4438).\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "ProxySender"
        }
      },
      {
        "value": "PubMedQA - Observed inference time (s)",
        "description": "PubMedQA is a biomedical question-answering dataset that evaluates a model's ability to interpret scientific literature. It consists of PubMed abstracts paired with yes/no/maybe questions derived from the content. The benchmark assesses a model's capability to reason over biomedical texts and provide factually grounded answers.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "PubMedQA"
        }
      },
      {
        "value": "EHRSQL - Observed inference time (s)",
        "description": "EHRSQL is a benchmark designed to evaluate models on generating structured queries for clinical research. Each example includes a natural language question and a database schema, and the task is to produce an SQL query that would return the correct result for a biomedical research objective. This benchmark assesses a model's understanding of medical terminology, data structures, and query construction.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "EHRSQL"
        }
      },
      {
        "value": "BMT-Status - Observed inference time (s)",
        "description": "BMT-Status is a benchmark composed of clinical notes and associated binary questions related to bone marrow transplant (BMT), hematopoietic stem cell transplant (HSCT), or hematopoietic cell transplant (HCT) status. The goal is to determine whether the patient received a subsequent transplant based on the provided clinical documentation.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "BMT-Status"
        }
      },
      {
        "value": "RaceBias - Observed inference time (s)",
        "description": "RaceBias is a benchmark used to evaluate language models for racially biased or inappropriate content in medical question-answering scenarios. Each instance consists of a medical question and a model-generated response. The task is to classify whether the response contains race-based, harmful, or inaccurate content. This benchmark supports research into bias detection and fairness in clinical AI systems.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "RaceBias"
        }
      },
      {
        "value": "N2C2-CT - Observed inference time (s)",
        "description": "A dataset that provides clinical notes and asks the model to classify whether the patient is a valid candidate for a provided clinical trial.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "N2C2-CT"
        }
      },
      {
        "value": "MedHallu - Observed inference time (s)",
        "description": "MedHallu is a benchmark focused on evaluating factual correctness in biomedical question answering. Each instance contains a PubMed-derived knowledge snippet, a biomedical question, and a model-generated answer. The task is to classify whether the answer is factually correct or contains hallucinated (non-grounded) information. This benchmark is designed to assess the factual reliability of medical language models.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MedHallu"
        }
      },
      {
        "value": "MIMIC-IV Billing Code - Observed inference time (s)",
        "description": "MIMIC-IV Billing Code is a benchmark derived from discharge summaries in the MIMIC-IV database, paired with their corresponding ICD-10 billing codes. The task requires models to extract structured billing codes based on free-text clinical notes, reflecting real-world hospital coding tasks for financial reimbursement.\n\nObserved inference runtime (s): Average observed time to process a request to the model (via an API, and thus depends on particular deployment).",
        "markdown": false,
        "lower_is_better": true,
        "metadata": {
          "metric": "Observed inference time (s)",
          "run_group": "MIMIC-IV Billing Code"
        }
      }
    ],
    "rows": [
      [
        {
          "value": "Claude 4.6 Opus",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.6086956521739131,
          "markdown": false
        },
        {
          "value": 3.9815040575374256,
          "description": "min=3.982, mean=3.982, max=3.982, sum=3.982 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 31.366606261746945,
          "description": "min=31.367, mean=31.367, max=31.367, sum=31.367 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 7.8456240898400695,
          "description": "min=7.846, mean=7.846, max=7.846, sum=7.846 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.3927846501397445,
          "description": "min=2.057, mean=2.393, max=3.318, sum=11.964 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=chemistry,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=medicine,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=pharmacology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=psychology,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 6.61648781191219,
          "description": "min=6.616, mean=6.616, max=6.616, sum=6.616 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.8352049187422734,
          "description": "min=2.835, mean=2.835, max=2.835, sum=2.835 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.887177960370506,
          "description": "min=2.887, mean=2.887, max=2.887, sum=2.887 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 15.450391584634781,
          "description": "min=15.45, mean=15.45, max=15.45, sum=15.45 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 25.718246800825,
          "description": "min=25.718, mean=25.718, max=25.718, sum=25.718 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4.579482823440121,
          "description": "min=4.579, mean=4.579, max=4.579, sum=4.579 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 7.40132422208786,
          "description": "min=7.401, mean=7.401, max=7.401, sum=7.401 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 10.376730933417775,
          "description": "min=10.377, mean=10.377, max=10.377, sum=10.377 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 30.477555736875534,
          "description": "min=30.478, mean=30.478, max=30.478, sum=30.478 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 3.35161071031606,
          "description": "min=3.309, mean=3.352, max=3.395, sum=6.703 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-opus-4.6",
            "med_dialog,subset=icliniq:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.233524876499176,
          "description": "min=2.234, mean=2.234, max=2.234, sum=2.234 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 13.532617599169413,
          "description": "min=13.533, mean=13.533, max=13.533, sum=13.533 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.151729763126373,
          "description": "min=2.152, mean=2.152, max=2.152, sum=2.152 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 3.0800699491500856,
          "description": "min=3.08, mean=3.08, max=3.08, sum=3.08 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.7476995937108746,
          "description": "min=2.748, mean=2.748, max=2.748, sum=2.748 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.139875107049942,
          "description": "min=2.14, mean=2.14, max=2.14, sum=2.14 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.587314378715561,
          "description": "min=2.587, mean=2.587, max=2.587, sum=2.587 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.2435376350145453,
          "description": "min=2.211, mean=2.244, max=2.275, sum=6.731 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2.8870265023708344,
          "description": "min=2.887, mean=2.887, max=2.887, sum=2.887 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Gemini 3.1 Pro (Preview)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.2217391304347826,
          "markdown": false
        },
        {
          "value": 21.206488949385555,
          "description": "min=21.206, mean=21.206, max=21.206, sum=21.206 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 21.93658061664054,
          "description": "min=21.937, mean=21.937, max=21.937, sum=21.937 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 37.10438707765422,
          "description": "min=37.104, mean=37.104, max=37.104, sum=37.104 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 6.889501849288109,
          "description": "min=5.505, mean=6.89, max=11.047, sum=34.448 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=medicine,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=psychology,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 20.24416685568822,
          "description": "min=20.244, mean=20.244, max=20.244, sum=20.244 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 8.09975773147589,
          "description": "min=8.1, mean=8.1, max=8.1, sum=8.1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 11.411923140223278,
          "description": "min=11.412, mean=11.412, max=11.412, sum=11.412 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 22.272048306465148,
          "description": "min=22.272, mean=22.272, max=22.272, sum=22.272 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 22.84699984639883,
          "description": "min=22.847, mean=22.847, max=22.847, sum=22.847 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 12.08153224640752,
          "description": "min=12.082, mean=12.082, max=12.082, sum=12.082 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 16.822261612415314,
          "description": "min=16.822, mean=16.822, max=16.822, sum=16.822 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 18.18655063144678,
          "description": "min=18.187, mean=18.187, max=18.187, sum=18.187 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 31.461765764045715,
          "description": "min=31.462, mean=31.462, max=31.462, sum=31.462 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 9.940946822639873,
          "description": "min=9.344, mean=9.941, max=10.538, sum=19.882 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.1-pro-preview",
            "med_dialog,subset=icliniq:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 7.539787130451202,
          "description": "min=7.54, mean=7.54, max=7.54, sum=7.54 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 25.243343569437663,
          "description": "min=25.243, mean=25.243, max=25.243, sum=25.243 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 12.368827115774154,
          "description": "min=12.369, mean=12.369, max=12.369, sum=12.369 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 8.885641211145991,
          "description": "min=8.886, mean=8.886, max=8.886, sum=8.886 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 28.814870612077918,
          "description": "min=28.815, mean=28.815, max=28.815, sum=28.815 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 10.106188394641876,
          "description": "min=10.106, mean=10.106, max=10.106, sum=10.106 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 10.264535181536646,
          "description": "min=10.265, mean=10.265, max=10.265, sum=10.265 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 7.820243248919571,
          "description": "min=6.363, mean=7.82, max=10.017, sum=23.461 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 7.705803503394127,
          "description": "min=7.706, mean=7.706, max=7.706, sum=7.706 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Gemini-3.5-flash",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.6434782608695652,
          "markdown": false
        },
        {
          "value": 7.703153204917908,
          "description": "min=7.703, mean=7.703, max=7.703, sum=7.703 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 10.30550671740494,
          "description": "min=10.306, mean=10.306, max=10.306, sum=10.306 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 11.471631377186608,
          "description": "min=11.472, mean=11.472, max=11.472, sum=11.472 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2.5126690640396254,
          "description": "min=2.068, mean=2.513, max=3.065, sum=12.563 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=medicine,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=psychology,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 5.533967949353255,
          "description": "min=5.534, mean=5.534, max=5.534, sum=5.534 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.3355365503705214,
          "description": "min=3.336, mean=3.336, max=3.336, sum=3.336 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.0158018529514874,
          "description": "min=3.016, mean=3.016, max=3.016, sum=3.016 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 11.022148422400157,
          "description": "min=11.022, mean=11.022, max=11.022, sum=11.022 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 11.201599702239037,
          "description": "min=11.202, mean=11.202, max=11.202, sum=11.202 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4.243845034094116,
          "description": "min=4.244, mean=4.244, max=4.244, sum=4.244 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 5.5809680008888245,
          "description": "min=5.581, mean=5.581, max=5.581, sum=5.581 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 9.415057977848026,
          "description": "min=9.415, mean=9.415, max=9.415, sum=9.415 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 12.983890476894379,
          "description": "min=12.984, mean=12.984, max=12.984, sum=12.984 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.829787428145611,
          "description": "min=3.828, mean=3.83, max=3.832, sum=7.66 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.5-flash",
            "med_dialog,subset=icliniq:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2.6625474193573,
          "description": "min=2.663, mean=2.663, max=2.663, sum=2.663 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 11.134681162834168,
          "description": "min=11.135, mean=11.135, max=11.135, sum=11.135 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4.1915532372951505,
          "description": "min=4.192, mean=4.192, max=4.192, sum=4.192 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.299833794116974,
          "description": "min=3.3, mean=3.3, max=3.3, sum=3.3 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 9.240120554757143,
          "description": "min=9.24, mean=9.24, max=9.24, sum=9.24 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2.949745077180862,
          "description": "min=2.95, mean=2.95, max=2.95, sum=2.95 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.9592316521855886,
          "description": "min=3.959, mean=3.959, max=3.959, sum=3.959 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2.3366370860695236,
          "description": "min=1.679, mean=2.337, max=3.458, sum=7.01 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 3.739136609630384,
          "description": "min=3.739, mean=3.739, max=3.739, sum=3.739 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.5-flash"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "GPT-5.4 (2026-03-05)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.3739130434782609,
          "markdown": false
        },
        {
          "value": 20.577027547501253,
          "description": "min=20.577, mean=20.577, max=20.577, sum=20.577 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 29.358331271580287,
          "description": "min=29.358, mean=29.358, max=29.358, sum=29.358 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 35.67437716418465,
          "description": "min=35.674, mean=35.674, max=35.674, sum=35.674 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 5.100587092844071,
          "description": "min=4.292, mean=5.101, max=6.092, sum=25.503 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 15.055612127502243,
          "description": "min=15.056, mean=15.056, max=15.056, sum=15.056 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 8.069247634322554,
          "description": "min=8.069, mean=8.069, max=8.069, sum=8.069 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 11.311257414762895,
          "description": "min=11.311, mean=11.311, max=11.311, sum=11.311 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 20.754599908987682,
          "description": "min=20.755, mean=20.755, max=20.755, sum=20.755 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 25.691876044496894,
          "description": "min=25.692, mean=25.692, max=25.692, sum=25.692 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 7.511928484595375,
          "description": "min=7.512, mean=7.512, max=7.512, sum=7.512 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 6.309806792736054,
          "description": "min=6.31, mean=6.31, max=6.31, sum=6.31 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 12.694286258633653,
          "description": "min=12.694, mean=12.694, max=12.694, sum=12.694 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 28.169949621755997,
          "description": "min=28.17, mean=28.17, max=28.17, sum=28.17 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.164852261793419,
          "description": "min=3.875, mean=4.165, max=4.454, sum=8.33 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 3.198045470237732,
          "description": "min=3.198, mean=3.198, max=3.198, sum=3.198 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 18.117119676272075,
          "description": "min=18.117, mean=18.117, max=18.117, sum=18.117 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 7.4045349000310265,
          "description": "min=7.405, mean=7.405, max=7.405, sum=7.405 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 5.675137461900711,
          "description": "min=5.675, mean=5.675, max=5.675, sum=5.675 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 16.339285349509854,
          "description": "min=16.339, mean=16.339, max=16.339, sum=16.339 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 3.929935823202133,
          "description": "min=3.93, mean=3.93, max=3.93, sum=3.93 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.155863125167207,
          "description": "min=4.156, mean=4.156, max=4.156, sum=4.156 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4.259084784864537,
          "description": "min=2.777, mean=4.259, max=6.339, sum=12.777 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 5.522191360153038,
          "description": "min=5.522, mean=5.522, max=5.522, sum=5.522 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "GPT-5.4 mini",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.6173913043478261,
          "markdown": false
        },
        {
          "value": 12.469016010327772,
          "description": "min=12.469, mean=12.469, max=12.469, sum=12.469 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 9.143483720841955,
          "description": "min=9.143, mean=9.143, max=9.143, sum=9.143 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 21.438040706535315,
          "description": "min=21.438, mean=21.438, max=21.438, sum=21.438 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.003209535035735,
          "description": "min=3.232, mean=4.003, max=4.683, sum=20.016 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 9.720431406776626,
          "description": "min=9.72, mean=9.72, max=9.72, sum=9.72 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 6.566571795462813,
          "description": "min=6.567, mean=6.567, max=6.567, sum=6.567 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 8.207329442203472,
          "description": "min=8.207, mean=8.207, max=8.207, sum=8.207 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 9.087990905841192,
          "description": "min=9.088, mean=9.088, max=9.088, sum=9.088 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 8.11744275316596,
          "description": "min=8.117, mean=8.117, max=8.117, sum=8.117 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.960133769026049,
          "description": "min=4.96, mean=4.96, max=4.96, sum=4.96 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.500222866535187,
          "description": "min=4.5, mean=4.5, max=4.5, sum=4.5 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 6.378590483158783,
          "description": "min=6.379, mean=6.379, max=6.379, sum=6.379 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 13.01923577092509,
          "description": "min=13.019, mean=13.019, max=13.019, sum=13.019 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 3.0771589608083207,
          "description": "min=2.969, mean=3.077, max=3.186, sum=6.154 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4-mini",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 2.567332151556015,
          "description": "min=2.567, mean=2.567, max=2.567, sum=2.567 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 7.220358753204346,
          "description": "min=7.22, mean=7.22, max=7.22, sum=7.22 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4.82353941411972,
          "description": "min=4.824, mean=4.824, max=4.824, sum=4.824 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 3.7834362304210662,
          "description": "min=3.783, mean=3.783, max=3.783, sum=3.783 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 14.35596393178361,
          "description": "min=14.356, mean=14.356, max=14.356, sum=14.356 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 2.994895841217041,
          "description": "min=2.995, mean=2.995, max=2.995, sum=2.995 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 2.7985065126133537,
          "description": "min=2.799, mean=2.799, max=2.799, sum=2.799 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 3.3213679672081633,
          "description": "min=2.74, mean=3.321, max=3.854, sum=9.964 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 5.62412586581707,
          "description": "min=5.624, mean=5.624, max=5.624, sum=5.624 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Claude 3.7 Sonnet (20250219)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.7347826086956522,
          "markdown": false
        },
        {
          "value": 3.863195901632309,
          "description": "min=3.863, mean=3.863, max=3.863, sum=3.863 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 13.848624661599723,
          "description": "min=13.849, mean=13.849, max=13.849, sum=13.849 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.4292833493582566,
          "description": "min=1.429, mean=1.429, max=1.429, sum=1.429 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.3447579393386841,
          "description": "min=1.345, mean=1.345, max=1.345, sum=1.345 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.252411273392764,
          "description": "min=1.252, mean=1.252, max=1.252, sum=1.252 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.6847498644417371,
          "description": "min=1.685, mean=1.685, max=1.685, sum=1.685 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.5195259516239166,
          "description": "min=1.52, mean=1.52, max=1.52, sum=1.52 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 13.21376874645551,
          "description": "min=13.214, mean=13.214, max=13.214, sum=13.214 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 12.927648959681392,
          "description": "min=12.928, mean=12.928, max=12.928, sum=12.928 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 3.2667968525886537,
          "description": "min=3.267, mean=3.267, max=3.267, sum=3.267 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 6.3847999048233035,
          "description": "min=6.385, mean=6.385, max=6.385, sum=6.385 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 8.33733369269461,
          "description": "min=8.337, mean=8.337, max=8.337, sum=8.337 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 8.463370934087484,
          "description": "min=8.463, mean=8.463, max=8.463, sum=8.463 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 3.1681859172582625,
          "description": "min=3.022, mean=3.168, max=3.314, sum=6.336 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "med_dialog,subset=icliniq:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 2.0782204871177674,
          "description": "min=2.078, mean=2.078, max=2.078, sum=2.078 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 11.665022039413453,
          "description": "min=11.665, mean=11.665, max=11.665, sum=11.665 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.765497265656789,
          "description": "min=1.765, mean=1.765, max=1.765, sum=1.765 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 2.4872676334381105,
          "description": "min=2.487, mean=2.487, max=2.487, sum=2.487 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 8.11839797091484,
          "description": "min=8.118, mean=8.118, max=8.118, sum=8.118 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 2.2075942852280357,
          "description": "min=2.208, mean=2.208, max=2.208, sum=2.208 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 2.5685647547601937,
          "description": "min=2.569, mean=2.569, max=2.569, sum=2.569 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 3.0745366978090867,
          "description": "min=3.06, mean=3.075, max=3.085, sum=9.224 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 4.534155822515488,
          "description": "min=4.534, mean=4.534, max=4.534, sum=4.534 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "DeepSeek R1",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.2608695652173913,
          "markdown": false
        },
        {
          "value": 43.75286227345467,
          "description": "min=43.753, mean=43.753, max=43.753, sum=43.753 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 29.774344934014714,
          "description": "min=29.774, mean=29.774, max=29.774, sum=29.774 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 41.87717197728117,
          "description": "min=41.877, mean=41.877, max=41.877, sum=41.877 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 20.78036990451813,
          "description": "min=20.78, mean=20.78, max=20.78, sum=20.78 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 34.088611539308125,
          "description": "min=34.089, mean=34.089, max=34.089, sum=34.089 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 10.915311019110487,
          "description": "min=10.915, mean=10.915, max=10.915, sum=10.915 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 8.023585789272042,
          "description": "min=8.024, mean=8.024, max=8.024, sum=8.024 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 15.778188250462215,
          "description": "min=15.778, mean=15.778, max=15.778, sum=15.778 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 21.84207220003009,
          "description": "min=21.842, mean=21.842, max=21.842, sum=21.842 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 12.422235173121615,
          "description": "min=12.422, mean=12.422, max=12.422, sum=12.422 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 8.098028691128047,
          "description": "min=8.098, mean=8.098, max=8.098, sum=8.098 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 21.51520863188521,
          "description": "min=21.515, mean=21.515, max=21.515, sum=21.515 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 21.370546659274115,
          "description": "min=21.371, mean=21.371, max=21.371, sum=21.371 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 8.633845594762542,
          "description": "min=7.447, mean=8.634, max=9.82, sum=17.268 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.726435486793518,
          "description": "min=4.726, mean=4.726, max=4.726, sum=4.726 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 21.951781096458436,
          "description": "min=21.952, mean=21.952, max=21.952, sum=21.952 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 5.393800741036733,
          "description": "min=5.394, mean=5.394, max=5.394, sum=5.394 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 8.807464688062668,
          "description": "min=8.807, mean=8.807, max=8.807, sum=8.807 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 19.193029651880263,
          "description": "min=19.193, mean=19.193, max=19.193, sum=19.193 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 4.999103357575157,
          "description": "min=4.999, mean=4.999, max=4.999, sum=4.999 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 7.5093938205056565,
          "description": "min=7.509, mean=7.509, max=7.509, sum=7.509 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 15.904167234127527,
          "description": "min=11.84, mean=15.904, max=23.456, sum=47.713 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 6.629145324707031,
          "description": "min=6.629, mean=6.629, max=6.629, sum=6.629 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Gemini 2.0 Flash",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.908695652173913,
          "markdown": false
        },
        {
          "value": 0.38494687390327453,
          "description": "min=0.385, mean=0.385, max=0.385, sum=0.385 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 5.227254615176198,
          "description": "min=5.227, mean=5.227, max=5.227, sum=5.227 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.4693439094664863,
          "description": "min=0.469, mean=0.469, max=0.469, sum=0.469 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.33297129392623903,
          "description": "min=0.333, mean=0.333, max=0.333, sum=0.333 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.33868198038695696,
          "description": "min=0.339, mean=0.339, max=0.339, sum=0.339 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1.7843808502136596,
          "description": "min=1.784, mean=1.784, max=1.784, sum=1.784 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1.7483746502399444,
          "description": "min=1.748, mean=1.748, max=1.748, sum=1.748 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 4.9915327548980715,
          "description": "min=4.992, mean=4.992, max=4.992, sum=4.992 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 5.109029924497008,
          "description": "min=5.109, mean=5.109, max=5.109, sum=5.109 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2.4756844749450684,
          "description": "min=2.476, mean=2.476, max=2.476, sum=2.476 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2.571005198955536,
          "description": "min=2.571, mean=2.571, max=2.571, sum=2.571 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 3.843229794190826,
          "description": "min=3.843, mean=3.843, max=3.843, sum=3.843 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 3.8159013593626154,
          "description": "min=3.816, mean=3.816, max=3.816, sum=3.816 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2.359554039001465,
          "description": "min=2.326, mean=2.36, max=2.393, sum=4.719 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "med_dialog,subset=icliniq:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2.023589251279831,
          "description": "min=2.024, mean=2.024, max=2.024, sum=2.024 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 5.9853372367223105,
          "description": "min=5.985, mean=5.985, max=5.985, sum=5.985 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2.35461799065272,
          "description": "min=2.355, mean=2.355, max=2.355, sum=2.355 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.3441569275856018,
          "description": "min=0.344, mean=0.344, max=0.344, sum=0.344 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.7652456421852112,
          "description": "min=0.765, mean=0.765, max=0.765, sum=0.765 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1.9761667815121737,
          "description": "min=1.976, mean=1.976, max=1.976, sum=1.976 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.40159591086610347,
          "description": "min=0.402, mean=0.402, max=0.402, sum=0.402 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2.5136772468108544,
          "description": "min=2.481, mean=2.514, max=2.551, sum=7.541 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.33655450963974,
          "description": "min=0.337, mean=0.337, max=0.337, sum=0.337 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Gemini 2.5 Pro (05-06 preview)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.1956521739130435,
          "markdown": false
        },
        {
          "value": 14.686965497255326,
          "description": "min=14.687, mean=14.687, max=14.687, sum=14.687 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 27.506609423099135,
          "description": "min=27.507, mean=27.507, max=27.507, sum=27.507 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 23.442691505853855,
          "description": "min=23.443, mean=23.443, max=23.443, sum=23.443 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 10.831548900842666,
          "description": "min=10.832, mean=10.832, max=10.832, sum=10.832 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 15.307000910306906,
          "description": "min=15.307, mean=15.307, max=15.307, sum=15.307 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 12.702731718836912,
          "description": "min=12.703, mean=12.703, max=12.703, sum=12.703 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 12.668443894386291,
          "description": "min=12.668, mean=12.668, max=12.668, sum=12.668 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 16.93871753414472,
          "description": "min=16.939, mean=16.939, max=16.939, sum=16.939 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 30.978275779634714,
          "description": "min=30.978, mean=30.978, max=30.978, sum=30.978 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 15.267056601285935,
          "description": "min=15.267, mean=15.267, max=15.267, sum=15.267 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 17.20419720649719,
          "description": "min=17.204, mean=17.204, max=17.204, sum=17.204 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 23.95976577545631,
          "description": "min=23.96, mean=23.96, max=23.96, sum=23.96 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 30.19600439269787,
          "description": "min=30.196, mean=30.196, max=30.196, sum=30.196 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 11.50131421494484,
          "description": "min=11.323, mean=11.501, max=11.68, sum=23.003 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 11.868826931715011,
          "description": "min=11.869, mean=11.869, max=11.869, sum=11.869 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 28.02089760462443,
          "description": "min=28.021, mean=28.021, max=28.021, sum=28.021 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 11.300854136943817,
          "description": "min=11.301, mean=11.301, max=11.301, sum=11.301 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 11.3364763879776,
          "description": "min=11.336, mean=11.336, max=11.336, sum=11.336 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 16.189389400970477,
          "description": "min=16.189, mean=16.189, max=16.189, sum=16.189 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 9.427392807873813,
          "description": "min=9.427, mean=9.427, max=9.427, sum=9.427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 12.19901739051956,
          "description": "min=12.199, mean=12.199, max=12.199, sum=12.199 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 12.830641206844831,
          "description": "min=8.357, mean=12.831, max=20.497, sum=38.492 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 12.977879849433899,
          "description": "min=12.978, mean=12.978, max=12.978, sum=12.978 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Llama 3.3 Instruct (70B)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.9260869565217391,
          "style": {
            "font-weight": "bold"
          },
          "markdown": false
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 9.148186227013083,
          "description": "min=9.148, mean=9.148, max=9.148, sum=9.148 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.24095447429877792,
          "description": "min=0.241, mean=0.241, max=0.241, sum=0.241 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.20930324862141886,
          "description": "min=0.209, mean=0.209, max=0.209, sum=0.209 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 11.603488028049469,
          "description": "min=11.603, mean=11.603, max=11.603, sum=11.603 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 9.823013886809349,
          "description": "min=9.823, mean=9.823, max=9.823, sum=9.823 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 5.954605188125219,
          "description": "min=5.955, mean=5.955, max=5.955, sum=5.955 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 2.7478469467163085,
          "description": "min=2.748, mean=2.748, max=2.748, sum=2.748 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 5.604794421631989,
          "description": "min=5.605, mean=5.605, max=5.605, sum=5.605 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 6.200638982397698,
          "description": "min=6.201, mean=6.201, max=6.201, sum=6.201 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.3253600368313856,
          "description": "min=1.291, mean=1.325, max=1.359, sum=2.651 (2)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "med_dialog,subset=icliniq:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.248099797487259,
          "description": "min=1.248, mean=1.248, max=1.248, sum=1.248 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 8.413016567230224,
          "description": "min=8.413, mean=8.413, max=8.413, sum=8.413 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.2374761176109314,
          "description": "min=0.237, mean=0.237, max=0.237, sum=0.237 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.2009574868462303,
          "description": "min=1.201, mean=1.201, max=1.201, sum=1.201 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.2071963594865427,
          "description": "min=1.117, mean=1.207, max=1.374, sum=3.622 (3)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=CREATININE,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {
            "font-weight": "bold"
          },
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Muse Spark (2026-04-08)",
          "description": "",
          "markdown": false
        },
        {
          "value": 0.008695652173913044,
          "markdown": false
        },
        {
          "value": 60.618242068507456,
          "description": "min=60.618, mean=60.618, max=60.618, sum=60.618 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 88.82792944707134,
          "description": "min=88.828, mean=88.828, max=88.828, sum=88.828 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 114.33869996102811,
          "description": "min=114.339, mean=114.339, max=114.339, sum=114.339 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 13.112066429113934,
          "description": "min=11.242, mean=13.112, max=17.498, sum=65.56 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=chemistry,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=medicine,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=pharmacology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=psychology,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 44.6799810877094,
          "description": "min=44.68, mean=44.68, max=44.68, sum=44.68 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 25.252749047687647,
          "description": "min=25.253, mean=25.253, max=25.253, sum=25.253 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 17.775139864055898,
          "description": "min=17.775, mean=17.775, max=17.775, sum=17.775 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 76.58062954545021,
          "description": "min=76.581, mean=76.581, max=76.581, sum=76.581 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 86.59814798086882,
          "description": "min=86.598, mean=86.598, max=86.598, sum=86.598 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 32.97713032095553,
          "description": "min=32.977, mean=32.977, max=32.977, sum=32.977 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 44.75749361038208,
          "description": "min=44.757, mean=44.757, max=44.757, sum=44.757 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 57.48885416361694,
          "description": "min=57.489, mean=57.489, max=57.489, sum=57.489 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 84.52363151779174,
          "description": "min=84.524, mean=84.524, max=84.524, sum=84.524 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 19.560871152852926,
          "description": "min=18.972, mean=19.561, max=20.15, sum=39.122 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_quiet_sand_v1alpha",
            "med_dialog,subset=icliniq:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 13.471088616085053,
          "description": "min=13.471, mean=13.471, max=13.471, sum=13.471 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 71.57455340544382,
          "description": "min=71.575, mean=71.575, max=71.575, sum=71.575 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 25.791178093385696,
          "description": "min=25.791, mean=25.791, max=25.791, sum=25.791 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 19.232449155569075,
          "description": "min=19.232, mean=19.232, max=19.232, sum=19.232 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 40.76429367040842,
          "description": "min=36.943, mean=40.764, max=44.586, sum=81.529 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 28.718790026044847,
          "description": "min=28.719, mean=28.719, max=28.719, sum=28.719 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 30.987977337694453,
          "description": "min=30.988, mean=30.988, max=30.988, sum=30.988 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 15.75721298506955,
          "description": "min=10.863, mean=15.757, max=24.744, sum=47.272 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=CREATININE,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 34.80593585085869,
          "description": "min=34.806, mean=34.806, max=34.806, sum=34.806 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ]
    ],
    "links": [
      {
        "text": "LaTeX",
        "href": "benchmarks/releases/v5.0.0/groups/latex/medhelm_scenarios_efficiency.tex"
      },
      {
        "text": "JSON",
        "href": "benchmarks/releases/v5.0.0/groups/json/medhelm_scenarios_efficiency.json"
      }
    ],
    "name": "efficiency"
  },
  {
    "title": "General information",
    "header": [
      {
        "value": "Model",
        "markdown": false,
        "metadata": {}
      },
      {
        "value": "MedCalc-Bench - # eval",
        "description": "MedCalc-Bench is a benchmark designed to evaluate models on their ability to compute clinically relevant values from patient notes. Each instance consists of a clinical note describing the patient's condition, a diagnostic question targeting a specific medical value, and a ground truth response. [(Khandekar et al., 2024)](https://arxiv.org/abs/2406.12036).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MedCalc-Bench"
        }
      },
      {
        "value": "MedCalc-Bench - # train",
        "description": "MedCalc-Bench is a benchmark designed to evaluate models on their ability to compute clinically relevant values from patient notes. Each instance consists of a clinical note describing the patient's condition, a diagnostic question targeting a specific medical value, and a ground truth response. [(Khandekar et al., 2024)](https://arxiv.org/abs/2406.12036).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MedCalc-Bench"
        }
      },
      {
        "value": "MedCalc-Bench - truncated",
        "description": "MedCalc-Bench is a benchmark designed to evaluate models on their ability to compute clinically relevant values from patient notes. Each instance consists of a clinical note describing the patient's condition, a diagnostic question targeting a specific medical value, and a ground truth response. [(Khandekar et al., 2024)](https://arxiv.org/abs/2406.12036).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MedCalc-Bench"
        }
      },
      {
        "value": "MedCalc-Bench - # prompt tokens",
        "description": "MedCalc-Bench is a benchmark designed to evaluate models on their ability to compute clinically relevant values from patient notes. Each instance consists of a clinical note describing the patient's condition, a diagnostic question targeting a specific medical value, and a ground truth response. [(Khandekar et al., 2024)](https://arxiv.org/abs/2406.12036).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MedCalc-Bench"
        }
      },
      {
        "value": "MedCalc-Bench - # output tokens",
        "description": "MedCalc-Bench is a benchmark designed to evaluate models on their ability to compute clinically relevant values from patient notes. Each instance consists of a clinical note describing the patient's condition, a diagnostic question targeting a specific medical value, and a ground truth response. [(Khandekar et al., 2024)](https://arxiv.org/abs/2406.12036).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MedCalc-Bench"
        }
      },
      {
        "value": "MTSamples - # eval",
        "description": "MTSamples Replicate is a benchmark that provides transcribed medical reports from various specialties. It is used to evaluate a model's ability to generate clinically appropriate treatment plans based on unstructured patient documentation [(MTSamples, 2025)](https://mtsamples.com).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MTSamples"
        }
      },
      {
        "value": "MTSamples - # train",
        "description": "MTSamples Replicate is a benchmark that provides transcribed medical reports from various specialties. It is used to evaluate a model's ability to generate clinically appropriate treatment plans based on unstructured patient documentation [(MTSamples, 2025)](https://mtsamples.com).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MTSamples"
        }
      },
      {
        "value": "MTSamples - truncated",
        "description": "MTSamples Replicate is a benchmark that provides transcribed medical reports from various specialties. It is used to evaluate a model's ability to generate clinically appropriate treatment plans based on unstructured patient documentation [(MTSamples, 2025)](https://mtsamples.com).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MTSamples"
        }
      },
      {
        "value": "MTSamples - # prompt tokens",
        "description": "MTSamples Replicate is a benchmark that provides transcribed medical reports from various specialties. It is used to evaluate a model's ability to generate clinically appropriate treatment plans based on unstructured patient documentation [(MTSamples, 2025)](https://mtsamples.com).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MTSamples"
        }
      },
      {
        "value": "MTSamples - # output tokens",
        "description": "MTSamples Replicate is a benchmark that provides transcribed medical reports from various specialties. It is used to evaluate a model's ability to generate clinically appropriate treatment plans based on unstructured patient documentation [(MTSamples, 2025)](https://mtsamples.com).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MTSamples"
        }
      },
      {
        "value": "Medec - # eval",
        "description": "Medec is a benchmark composed of clinical narratives that include either correct documentation or medical errors. Each entry includes sentence-level identifiers and an associated correction task. The model must review the narrative and either identify the erroneous sentence and correct it, or confirm that the text is entirely accurate [(Abacha et al., 2025)](https://arxiv.org/abs/2412.19260).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "Medec"
        }
      },
      {
        "value": "Medec - # train",
        "description": "Medec is a benchmark composed of clinical narratives that include either correct documentation or medical errors. Each entry includes sentence-level identifiers and an associated correction task. The model must review the narrative and either identify the erroneous sentence and correct it, or confirm that the text is entirely accurate [(Abacha et al., 2025)](https://arxiv.org/abs/2412.19260).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "Medec"
        }
      },
      {
        "value": "Medec - truncated",
        "description": "Medec is a benchmark composed of clinical narratives that include either correct documentation or medical errors. Each entry includes sentence-level identifiers and an associated correction task. The model must review the narrative and either identify the erroneous sentence and correct it, or confirm that the text is entirely accurate [(Abacha et al., 2025)](https://arxiv.org/abs/2412.19260).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "Medec"
        }
      },
      {
        "value": "Medec - # prompt tokens",
        "description": "Medec is a benchmark composed of clinical narratives that include either correct documentation or medical errors. Each entry includes sentence-level identifiers and an associated correction task. The model must review the narrative and either identify the erroneous sentence and correct it, or confirm that the text is entirely accurate [(Abacha et al., 2025)](https://arxiv.org/abs/2412.19260).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "Medec"
        }
      },
      {
        "value": "Medec - # output tokens",
        "description": "Medec is a benchmark composed of clinical narratives that include either correct documentation or medical errors. Each entry includes sentence-level identifiers and an associated correction task. The model must review the narrative and either identify the erroneous sentence and correct it, or confirm that the text is entirely accurate [(Abacha et al., 2025)](https://arxiv.org/abs/2412.19260).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "Medec"
        }
      },
      {
        "value": "HeadQA - # eval",
        "description": "HeadQA is a benchmark consisting of biomedical multiple-choice questions intended to evaluate a model's medical knowledge and reasoning. Each instance presents a clinical or scientific question with four answer options, requiring the model to select the most appropriate answer [(Vilares et al., 2019)](https://arxiv.org/abs/1906.04701).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "HeadQA"
        }
      },
      {
        "value": "HeadQA - # train",
        "description": "HeadQA is a benchmark consisting of biomedical multiple-choice questions intended to evaluate a model's medical knowledge and reasoning. Each instance presents a clinical or scientific question with four answer options, requiring the model to select the most appropriate answer [(Vilares et al., 2019)](https://arxiv.org/abs/1906.04701).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "HeadQA"
        }
      },
      {
        "value": "HeadQA - truncated",
        "description": "HeadQA is a benchmark consisting of biomedical multiple-choice questions intended to evaluate a model's medical knowledge and reasoning. Each instance presents a clinical or scientific question with four answer options, requiring the model to select the most appropriate answer [(Vilares et al., 2019)](https://arxiv.org/abs/1906.04701).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "HeadQA"
        }
      },
      {
        "value": "HeadQA - # prompt tokens",
        "description": "HeadQA is a benchmark consisting of biomedical multiple-choice questions intended to evaluate a model's medical knowledge and reasoning. Each instance presents a clinical or scientific question with four answer options, requiring the model to select the most appropriate answer [(Vilares et al., 2019)](https://arxiv.org/abs/1906.04701).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "HeadQA"
        }
      },
      {
        "value": "HeadQA - # output tokens",
        "description": "HeadQA is a benchmark consisting of biomedical multiple-choice questions intended to evaluate a model's medical knowledge and reasoning. Each instance presents a clinical or scientific question with four answer options, requiring the model to select the most appropriate answer [(Vilares et al., 2019)](https://arxiv.org/abs/1906.04701).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "HeadQA"
        }
      },
      {
        "value": "Medbullets - # eval",
        "description": "Medbullets is a benchmark of USMLE-style medical questions designed to assess a model's ability to understand and apply clinical knowledge. Each question is accompanied by a patient scenario and five multiple-choice options, similar to those found on Step 2 and Step 3 board exams [(MedBullets, 2025)](https://step2.medbullets.com).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "Medbullets"
        }
      },
      {
        "value": "Medbullets - # train",
        "description": "Medbullets is a benchmark of USMLE-style medical questions designed to assess a model's ability to understand and apply clinical knowledge. Each question is accompanied by a patient scenario and five multiple-choice options, similar to those found on Step 2 and Step 3 board exams [(MedBullets, 2025)](https://step2.medbullets.com).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "Medbullets"
        }
      },
      {
        "value": "Medbullets - truncated",
        "description": "Medbullets is a benchmark of USMLE-style medical questions designed to assess a model's ability to understand and apply clinical knowledge. Each question is accompanied by a patient scenario and five multiple-choice options, similar to those found on Step 2 and Step 3 board exams [(MedBullets, 2025)](https://step2.medbullets.com).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "Medbullets"
        }
      },
      {
        "value": "Medbullets - # prompt tokens",
        "description": "Medbullets is a benchmark of USMLE-style medical questions designed to assess a model's ability to understand and apply clinical knowledge. Each question is accompanied by a patient scenario and five multiple-choice options, similar to those found on Step 2 and Step 3 board exams [(MedBullets, 2025)](https://step2.medbullets.com).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "Medbullets"
        }
      },
      {
        "value": "Medbullets - # output tokens",
        "description": "Medbullets is a benchmark of USMLE-style medical questions designed to assess a model's ability to understand and apply clinical knowledge. Each question is accompanied by a patient scenario and five multiple-choice options, similar to those found on Step 2 and Step 3 board exams [(MedBullets, 2025)](https://step2.medbullets.com).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "Medbullets"
        }
      },
      {
        "value": "MedQA - # eval",
        "description": "MedQA is an open domain question answering dataset composed of questions from professional medical board exams ([Jin et al. 2020](https://arxiv.org/pdf/2009.13081.pdf)).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MedQA"
        }
      },
      {
        "value": "MedQA - # train",
        "description": "MedQA is an open domain question answering dataset composed of questions from professional medical board exams ([Jin et al. 2020](https://arxiv.org/pdf/2009.13081.pdf)).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MedQA"
        }
      },
      {
        "value": "MedQA - truncated",
        "description": "MedQA is an open domain question answering dataset composed of questions from professional medical board exams ([Jin et al. 2020](https://arxiv.org/pdf/2009.13081.pdf)).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MedQA"
        }
      },
      {
        "value": "MedQA - # prompt tokens",
        "description": "MedQA is an open domain question answering dataset composed of questions from professional medical board exams ([Jin et al. 2020](https://arxiv.org/pdf/2009.13081.pdf)).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MedQA"
        }
      },
      {
        "value": "MedQA - # output tokens",
        "description": "MedQA is an open domain question answering dataset composed of questions from professional medical board exams ([Jin et al. 2020](https://arxiv.org/pdf/2009.13081.pdf)).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MedQA"
        }
      },
      {
        "value": "MedMCQA - # eval",
        "description": "MedMCQA is a multiple-choice question answering (MCQA) dataset designed to address real-world medical entrance exam questions ([Flores et al. 2020](https://arxiv.org/abs/2203.14371)).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MedMCQA"
        }
      },
      {
        "value": "MedMCQA - # train",
        "description": "MedMCQA is a multiple-choice question answering (MCQA) dataset designed to address real-world medical entrance exam questions ([Flores et al. 2020](https://arxiv.org/abs/2203.14371)).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MedMCQA"
        }
      },
      {
        "value": "MedMCQA - truncated",
        "description": "MedMCQA is a multiple-choice question answering (MCQA) dataset designed to address real-world medical entrance exam questions ([Flores et al. 2020](https://arxiv.org/abs/2203.14371)).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MedMCQA"
        }
      },
      {
        "value": "MedMCQA - # prompt tokens",
        "description": "MedMCQA is a multiple-choice question answering (MCQA) dataset designed to address real-world medical entrance exam questions ([Flores et al. 2020](https://arxiv.org/abs/2203.14371)).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MedMCQA"
        }
      },
      {
        "value": "MedMCQA - # output tokens",
        "description": "MedMCQA is a multiple-choice question answering (MCQA) dataset designed to address real-world medical entrance exam questions ([Flores et al. 2020](https://arxiv.org/abs/2203.14371)).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MedMCQA"
        }
      },
      {
        "value": "ACI-Bench - # eval",
        "description": "ACI-Bench is a benchmark of real-world patient-doctor conversations paired with structured clinical notes. The benchmark evaluates a model's ability to understand spoken medical dialogue and convert it into formal clinical documentation, covering sections such as history of present illness, physical exam findings, results, and assessment and plan [(Yim et al., 2024)](https://www.nature.com/articles/s41597-023-02487-3).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "ACI-Bench"
        }
      },
      {
        "value": "ACI-Bench - # train",
        "description": "ACI-Bench is a benchmark of real-world patient-doctor conversations paired with structured clinical notes. The benchmark evaluates a model's ability to understand spoken medical dialogue and convert it into formal clinical documentation, covering sections such as history of present illness, physical exam findings, results, and assessment and plan [(Yim et al., 2024)](https://www.nature.com/articles/s41597-023-02487-3).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "ACI-Bench"
        }
      },
      {
        "value": "ACI-Bench - truncated",
        "description": "ACI-Bench is a benchmark of real-world patient-doctor conversations paired with structured clinical notes. The benchmark evaluates a model's ability to understand spoken medical dialogue and convert it into formal clinical documentation, covering sections such as history of present illness, physical exam findings, results, and assessment and plan [(Yim et al., 2024)](https://www.nature.com/articles/s41597-023-02487-3).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "ACI-Bench"
        }
      },
      {
        "value": "ACI-Bench - # prompt tokens",
        "description": "ACI-Bench is a benchmark of real-world patient-doctor conversations paired with structured clinical notes. The benchmark evaluates a model's ability to understand spoken medical dialogue and convert it into formal clinical documentation, covering sections such as history of present illness, physical exam findings, results, and assessment and plan [(Yim et al., 2024)](https://www.nature.com/articles/s41597-023-02487-3).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "ACI-Bench"
        }
      },
      {
        "value": "ACI-Bench - # output tokens",
        "description": "ACI-Bench is a benchmark of real-world patient-doctor conversations paired with structured clinical notes. The benchmark evaluates a model's ability to understand spoken medical dialogue and convert it into formal clinical documentation, covering sections such as history of present illness, physical exam findings, results, and assessment and plan [(Yim et al., 2024)](https://www.nature.com/articles/s41597-023-02487-3).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "ACI-Bench"
        }
      },
      {
        "value": "MTSamples Procedures - # eval",
        "description": "MTSamples Procedures is a benchmark composed of transcribed operative notes, focused on documenting surgical procedures. Each example presents a brief patient case involving a surgical intervention, and the model is tasked with generating a coherent and clinically accurate procedural summary or treatment plan.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MTSamples Procedures"
        }
      },
      {
        "value": "MTSamples Procedures - # train",
        "description": "MTSamples Procedures is a benchmark composed of transcribed operative notes, focused on documenting surgical procedures. Each example presents a brief patient case involving a surgical intervention, and the model is tasked with generating a coherent and clinically accurate procedural summary or treatment plan.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MTSamples Procedures"
        }
      },
      {
        "value": "MTSamples Procedures - truncated",
        "description": "MTSamples Procedures is a benchmark composed of transcribed operative notes, focused on documenting surgical procedures. Each example presents a brief patient case involving a surgical intervention, and the model is tasked with generating a coherent and clinically accurate procedural summary or treatment plan.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MTSamples Procedures"
        }
      },
      {
        "value": "MTSamples Procedures - # prompt tokens",
        "description": "MTSamples Procedures is a benchmark composed of transcribed operative notes, focused on documenting surgical procedures. Each example presents a brief patient case involving a surgical intervention, and the model is tasked with generating a coherent and clinically accurate procedural summary or treatment plan.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MTSamples Procedures"
        }
      },
      {
        "value": "MTSamples Procedures - # output tokens",
        "description": "MTSamples Procedures is a benchmark composed of transcribed operative notes, focused on documenting surgical procedures. Each example presents a brief patient case involving a surgical intervention, and the model is tasked with generating a coherent and clinically accurate procedural summary or treatment plan.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MTSamples Procedures"
        }
      },
      {
        "value": "MIMIC-RRS - # eval",
        "description": "MIMIC-RRS is a benchmark constructed from radiology reports in the MIMIC-III database. It contains pairs of \u2018Findings\u2018 and \u2018Impression\u2018 sections, enabling evaluation of a model's ability to summarize diagnostic imaging observations into concise, clinically relevant conclusions [(Chen et al., 2023)](https://arxiv.org/abs/2211.08584).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MIMIC-RRS"
        }
      },
      {
        "value": "MIMIC-RRS - # train",
        "description": "MIMIC-RRS is a benchmark constructed from radiology reports in the MIMIC-III database. It contains pairs of \u2018Findings\u2018 and \u2018Impression\u2018 sections, enabling evaluation of a model's ability to summarize diagnostic imaging observations into concise, clinically relevant conclusions [(Chen et al., 2023)](https://arxiv.org/abs/2211.08584).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MIMIC-RRS"
        }
      },
      {
        "value": "MIMIC-RRS - truncated",
        "description": "MIMIC-RRS is a benchmark constructed from radiology reports in the MIMIC-III database. It contains pairs of \u2018Findings\u2018 and \u2018Impression\u2018 sections, enabling evaluation of a model's ability to summarize diagnostic imaging observations into concise, clinically relevant conclusions [(Chen et al., 2023)](https://arxiv.org/abs/2211.08584).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MIMIC-RRS"
        }
      },
      {
        "value": "MIMIC-RRS - # prompt tokens",
        "description": "MIMIC-RRS is a benchmark constructed from radiology reports in the MIMIC-III database. It contains pairs of \u2018Findings\u2018 and \u2018Impression\u2018 sections, enabling evaluation of a model's ability to summarize diagnostic imaging observations into concise, clinically relevant conclusions [(Chen et al., 2023)](https://arxiv.org/abs/2211.08584).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MIMIC-RRS"
        }
      },
      {
        "value": "MIMIC-RRS - # output tokens",
        "description": "MIMIC-RRS is a benchmark constructed from radiology reports in the MIMIC-III database. It contains pairs of \u2018Findings\u2018 and \u2018Impression\u2018 sections, enabling evaluation of a model's ability to summarize diagnostic imaging observations into concise, clinically relevant conclusions [(Chen et al., 2023)](https://arxiv.org/abs/2211.08584).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MIMIC-RRS"
        }
      },
      {
        "value": "MIMIC-BHC - # eval",
        "description": "MIMIC-BHC is a benchmark focused on summarization of discharge notes into Brief Hospital Course (BHC) sections. It consists of curated discharge notes from MIMIC-IV, each paired with its corresponding BHC summary. The benchmark evaluates a model's ability to condense detailed clinical information into accurate, concise summaries that reflect the patient's hospital stay [(Aali et al., 2024)](https://doi.org/10.1093/jamia/ocae312).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MIMIC-BHC"
        }
      },
      {
        "value": "MIMIC-BHC - # train",
        "description": "MIMIC-BHC is a benchmark focused on summarization of discharge notes into Brief Hospital Course (BHC) sections. It consists of curated discharge notes from MIMIC-IV, each paired with its corresponding BHC summary. The benchmark evaluates a model's ability to condense detailed clinical information into accurate, concise summaries that reflect the patient's hospital stay [(Aali et al., 2024)](https://doi.org/10.1093/jamia/ocae312).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MIMIC-BHC"
        }
      },
      {
        "value": "MIMIC-BHC - truncated",
        "description": "MIMIC-BHC is a benchmark focused on summarization of discharge notes into Brief Hospital Course (BHC) sections. It consists of curated discharge notes from MIMIC-IV, each paired with its corresponding BHC summary. The benchmark evaluates a model's ability to condense detailed clinical information into accurate, concise summaries that reflect the patient's hospital stay [(Aali et al., 2024)](https://doi.org/10.1093/jamia/ocae312).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MIMIC-BHC"
        }
      },
      {
        "value": "MIMIC-BHC - # prompt tokens",
        "description": "MIMIC-BHC is a benchmark focused on summarization of discharge notes into Brief Hospital Course (BHC) sections. It consists of curated discharge notes from MIMIC-IV, each paired with its corresponding BHC summary. The benchmark evaluates a model's ability to condense detailed clinical information into accurate, concise summaries that reflect the patient's hospital stay [(Aali et al., 2024)](https://doi.org/10.1093/jamia/ocae312).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MIMIC-BHC"
        }
      },
      {
        "value": "MIMIC-BHC - # output tokens",
        "description": "MIMIC-BHC is a benchmark focused on summarization of discharge notes into Brief Hospital Course (BHC) sections. It consists of curated discharge notes from MIMIC-IV, each paired with its corresponding BHC summary. The benchmark evaluates a model's ability to condense detailed clinical information into accurate, concise summaries that reflect the patient's hospital stay [(Aali et al., 2024)](https://doi.org/10.1093/jamia/ocae312).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MIMIC-BHC"
        }
      },
      {
        "value": "MedicationQA - # eval",
        "description": "MedicationQA is a benchmark composed of open-ended consumer health questions specifically focused on medications. Each example consists of a free-form question and a corresponding medically grounded answer. The benchmark evaluates a model's ability to provide accurate, accessible, and informative medication-related responses for a lay audience.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MedicationQA"
        }
      },
      {
        "value": "MedicationQA - # train",
        "description": "MedicationQA is a benchmark composed of open-ended consumer health questions specifically focused on medications. Each example consists of a free-form question and a corresponding medically grounded answer. The benchmark evaluates a model's ability to provide accurate, accessible, and informative medication-related responses for a lay audience.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MedicationQA"
        }
      },
      {
        "value": "MedicationQA - truncated",
        "description": "MedicationQA is a benchmark composed of open-ended consumer health questions specifically focused on medications. Each example consists of a free-form question and a corresponding medically grounded answer. The benchmark evaluates a model's ability to provide accurate, accessible, and informative medication-related responses for a lay audience.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MedicationQA"
        }
      },
      {
        "value": "MedicationQA - # prompt tokens",
        "description": "MedicationQA is a benchmark composed of open-ended consumer health questions specifically focused on medications. Each example consists of a free-form question and a corresponding medically grounded answer. The benchmark evaluates a model's ability to provide accurate, accessible, and informative medication-related responses for a lay audience.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MedicationQA"
        }
      },
      {
        "value": "MedicationQA - # output tokens",
        "description": "MedicationQA is a benchmark composed of open-ended consumer health questions specifically focused on medications. Each example consists of a free-form question and a corresponding medically grounded answer. The benchmark evaluates a model's ability to provide accurate, accessible, and informative medication-related responses for a lay audience.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MedicationQA"
        }
      },
      {
        "value": "PatientInstruct - # eval",
        "description": "PatientInstruct is a benchmark designed to evaluate models on generating personalized post-procedure instructions for patients. It includes real-world clinical case details, such as diagnosis, planned procedures, and history and physical notes, from which models must produce clear, actionable instructions appropriate for patients recovering from medical interventions.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "PatientInstruct"
        }
      },
      {
        "value": "PatientInstruct - # train",
        "description": "PatientInstruct is a benchmark designed to evaluate models on generating personalized post-procedure instructions for patients. It includes real-world clinical case details, such as diagnosis, planned procedures, and history and physical notes, from which models must produce clear, actionable instructions appropriate for patients recovering from medical interventions.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "PatientInstruct"
        }
      },
      {
        "value": "PatientInstruct - truncated",
        "description": "PatientInstruct is a benchmark designed to evaluate models on generating personalized post-procedure instructions for patients. It includes real-world clinical case details, such as diagnosis, planned procedures, and history and physical notes, from which models must produce clear, actionable instructions appropriate for patients recovering from medical interventions.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "PatientInstruct"
        }
      },
      {
        "value": "PatientInstruct - # prompt tokens",
        "description": "PatientInstruct is a benchmark designed to evaluate models on generating personalized post-procedure instructions for patients. It includes real-world clinical case details, such as diagnosis, planned procedures, and history and physical notes, from which models must produce clear, actionable instructions appropriate for patients recovering from medical interventions.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "PatientInstruct"
        }
      },
      {
        "value": "PatientInstruct - # output tokens",
        "description": "PatientInstruct is a benchmark designed to evaluate models on generating personalized post-procedure instructions for patients. It includes real-world clinical case details, such as diagnosis, planned procedures, and history and physical notes, from which models must produce clear, actionable instructions appropriate for patients recovering from medical interventions.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "PatientInstruct"
        }
      },
      {
        "value": "MedDialog - # eval",
        "description": "MedDialog is a benchmark of real-world doctor-patient conversations focused on health-related concerns and advice. Each dialogue is paired with a one-sentence summary that reflects the core patient question or exchange. The benchmark evaluates a model's ability to condense medical dialogue into concise, informative summaries.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MedDialog"
        }
      },
      {
        "value": "MedDialog - # train",
        "description": "MedDialog is a benchmark of real-world doctor-patient conversations focused on health-related concerns and advice. Each dialogue is paired with a one-sentence summary that reflects the core patient question or exchange. The benchmark evaluates a model's ability to condense medical dialogue into concise, informative summaries.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MedDialog"
        }
      },
      {
        "value": "MedDialog - truncated",
        "description": "MedDialog is a benchmark of real-world doctor-patient conversations focused on health-related concerns and advice. Each dialogue is paired with a one-sentence summary that reflects the core patient question or exchange. The benchmark evaluates a model's ability to condense medical dialogue into concise, informative summaries.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MedDialog"
        }
      },
      {
        "value": "MedDialog - # prompt tokens",
        "description": "MedDialog is a benchmark of real-world doctor-patient conversations focused on health-related concerns and advice. Each dialogue is paired with a one-sentence summary that reflects the core patient question or exchange. The benchmark evaluates a model's ability to condense medical dialogue into concise, informative summaries.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MedDialog"
        }
      },
      {
        "value": "MedDialog - # output tokens",
        "description": "MedDialog is a benchmark of real-world doctor-patient conversations focused on health-related concerns and advice. Each dialogue is paired with a one-sentence summary that reflects the core patient question or exchange. The benchmark evaluates a model's ability to condense medical dialogue into concise, informative summaries.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MedDialog"
        }
      },
      {
        "value": "MedConfInfo - # eval",
        "description": "MedConfInfo is a benchmark comprising clinical notes from adolescent patients. It is used to evaluate whether the content contains sensitive protected health information (PHI) that should be restricted from parental access, in accordance with adolescent confidentiality policies in clinical care. [(Rabbani et al., 2024)](https://jamanetwork.com/journals/jamapediatrics/fullarticle/2814109).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MedConfInfo"
        }
      },
      {
        "value": "MedConfInfo - # train",
        "description": "MedConfInfo is a benchmark comprising clinical notes from adolescent patients. It is used to evaluate whether the content contains sensitive protected health information (PHI) that should be restricted from parental access, in accordance with adolescent confidentiality policies in clinical care. [(Rabbani et al., 2024)](https://jamanetwork.com/journals/jamapediatrics/fullarticle/2814109).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MedConfInfo"
        }
      },
      {
        "value": "MedConfInfo - truncated",
        "description": "MedConfInfo is a benchmark comprising clinical notes from adolescent patients. It is used to evaluate whether the content contains sensitive protected health information (PHI) that should be restricted from parental access, in accordance with adolescent confidentiality policies in clinical care. [(Rabbani et al., 2024)](https://jamanetwork.com/journals/jamapediatrics/fullarticle/2814109).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MedConfInfo"
        }
      },
      {
        "value": "MedConfInfo - # prompt tokens",
        "description": "MedConfInfo is a benchmark comprising clinical notes from adolescent patients. It is used to evaluate whether the content contains sensitive protected health information (PHI) that should be restricted from parental access, in accordance with adolescent confidentiality policies in clinical care. [(Rabbani et al., 2024)](https://jamanetwork.com/journals/jamapediatrics/fullarticle/2814109).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MedConfInfo"
        }
      },
      {
        "value": "MedConfInfo - # output tokens",
        "description": "MedConfInfo is a benchmark comprising clinical notes from adolescent patients. It is used to evaluate whether the content contains sensitive protected health information (PHI) that should be restricted from parental access, in accordance with adolescent confidentiality policies in clinical care. [(Rabbani et al., 2024)](https://jamanetwork.com/journals/jamapediatrics/fullarticle/2814109).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MedConfInfo"
        }
      },
      {
        "value": "MEDIQA - # eval",
        "description": "MEDIQA is a benchmark designed to evaluate a model's ability to retrieve and generate medically accurate answers to patient-generated questions. Each instance includes a consumer health question, a set of candidate answers (used in ranking tasks), relevance annotations, and optionally, additional context. The benchmark focuses on supporting patient understanding and accessibility in health communication.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MEDIQA"
        }
      },
      {
        "value": "MEDIQA - # train",
        "description": "MEDIQA is a benchmark designed to evaluate a model's ability to retrieve and generate medically accurate answers to patient-generated questions. Each instance includes a consumer health question, a set of candidate answers (used in ranking tasks), relevance annotations, and optionally, additional context. The benchmark focuses on supporting patient understanding and accessibility in health communication.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MEDIQA"
        }
      },
      {
        "value": "MEDIQA - truncated",
        "description": "MEDIQA is a benchmark designed to evaluate a model's ability to retrieve and generate medically accurate answers to patient-generated questions. Each instance includes a consumer health question, a set of candidate answers (used in ranking tasks), relevance annotations, and optionally, additional context. The benchmark focuses on supporting patient understanding and accessibility in health communication.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MEDIQA"
        }
      },
      {
        "value": "MEDIQA - # prompt tokens",
        "description": "MEDIQA is a benchmark designed to evaluate a model's ability to retrieve and generate medically accurate answers to patient-generated questions. Each instance includes a consumer health question, a set of candidate answers (used in ranking tasks), relevance annotations, and optionally, additional context. The benchmark focuses on supporting patient understanding and accessibility in health communication.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MEDIQA"
        }
      },
      {
        "value": "MEDIQA - # output tokens",
        "description": "MEDIQA is a benchmark designed to evaluate a model's ability to retrieve and generate medically accurate answers to patient-generated questions. Each instance includes a consumer health question, a set of candidate answers (used in ranking tasks), relevance annotations, and optionally, additional context. The benchmark focuses on supporting patient understanding and accessibility in health communication.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MEDIQA"
        }
      },
      {
        "value": "ProxySender - # eval",
        "description": "ProxySender is a benchmark composed of patient portal messages received by clinicians. It evaluates whether the message was sent by the patient or by a proxy user (e.g., parent, spouse), which is critical for understanding who is communicating with healthcare providers. [(Tse G, et al., 2025)](https://doi.org/10.1001/jamapediatrics.2024.4438).\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "ProxySender"
        }
      },
      {
        "value": "ProxySender - # train",
        "description": "ProxySender is a benchmark composed of patient portal messages received by clinicians. It evaluates whether the message was sent by the patient or by a proxy user (e.g., parent, spouse), which is critical for understanding who is communicating with healthcare providers. [(Tse G, et al., 2025)](https://doi.org/10.1001/jamapediatrics.2024.4438).\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "ProxySender"
        }
      },
      {
        "value": "ProxySender - truncated",
        "description": "ProxySender is a benchmark composed of patient portal messages received by clinicians. It evaluates whether the message was sent by the patient or by a proxy user (e.g., parent, spouse), which is critical for understanding who is communicating with healthcare providers. [(Tse G, et al., 2025)](https://doi.org/10.1001/jamapediatrics.2024.4438).\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "ProxySender"
        }
      },
      {
        "value": "ProxySender - # prompt tokens",
        "description": "ProxySender is a benchmark composed of patient portal messages received by clinicians. It evaluates whether the message was sent by the patient or by a proxy user (e.g., parent, spouse), which is critical for understanding who is communicating with healthcare providers. [(Tse G, et al., 2025)](https://doi.org/10.1001/jamapediatrics.2024.4438).\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "ProxySender"
        }
      },
      {
        "value": "ProxySender - # output tokens",
        "description": "ProxySender is a benchmark composed of patient portal messages received by clinicians. It evaluates whether the message was sent by the patient or by a proxy user (e.g., parent, spouse), which is critical for understanding who is communicating with healthcare providers. [(Tse G, et al., 2025)](https://doi.org/10.1001/jamapediatrics.2024.4438).\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "ProxySender"
        }
      },
      {
        "value": "PubMedQA - # eval",
        "description": "PubMedQA is a biomedical question-answering dataset that evaluates a model's ability to interpret scientific literature. It consists of PubMed abstracts paired with yes/no/maybe questions derived from the content. The benchmark assesses a model's capability to reason over biomedical texts and provide factually grounded answers.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "PubMedQA"
        }
      },
      {
        "value": "PubMedQA - # train",
        "description": "PubMedQA is a biomedical question-answering dataset that evaluates a model's ability to interpret scientific literature. It consists of PubMed abstracts paired with yes/no/maybe questions derived from the content. The benchmark assesses a model's capability to reason over biomedical texts and provide factually grounded answers.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "PubMedQA"
        }
      },
      {
        "value": "PubMedQA - truncated",
        "description": "PubMedQA is a biomedical question-answering dataset that evaluates a model's ability to interpret scientific literature. It consists of PubMed abstracts paired with yes/no/maybe questions derived from the content. The benchmark assesses a model's capability to reason over biomedical texts and provide factually grounded answers.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "PubMedQA"
        }
      },
      {
        "value": "PubMedQA - # prompt tokens",
        "description": "PubMedQA is a biomedical question-answering dataset that evaluates a model's ability to interpret scientific literature. It consists of PubMed abstracts paired with yes/no/maybe questions derived from the content. The benchmark assesses a model's capability to reason over biomedical texts and provide factually grounded answers.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "PubMedQA"
        }
      },
      {
        "value": "PubMedQA - # output tokens",
        "description": "PubMedQA is a biomedical question-answering dataset that evaluates a model's ability to interpret scientific literature. It consists of PubMed abstracts paired with yes/no/maybe questions derived from the content. The benchmark assesses a model's capability to reason over biomedical texts and provide factually grounded answers.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "PubMedQA"
        }
      },
      {
        "value": "EHRSQL - # eval",
        "description": "EHRSQL is a benchmark designed to evaluate models on generating structured queries for clinical research. Each example includes a natural language question and a database schema, and the task is to produce an SQL query that would return the correct result for a biomedical research objective. This benchmark assesses a model's understanding of medical terminology, data structures, and query construction.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "EHRSQL"
        }
      },
      {
        "value": "EHRSQL - # train",
        "description": "EHRSQL is a benchmark designed to evaluate models on generating structured queries for clinical research. Each example includes a natural language question and a database schema, and the task is to produce an SQL query that would return the correct result for a biomedical research objective. This benchmark assesses a model's understanding of medical terminology, data structures, and query construction.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "EHRSQL"
        }
      },
      {
        "value": "EHRSQL - truncated",
        "description": "EHRSQL is a benchmark designed to evaluate models on generating structured queries for clinical research. Each example includes a natural language question and a database schema, and the task is to produce an SQL query that would return the correct result for a biomedical research objective. This benchmark assesses a model's understanding of medical terminology, data structures, and query construction.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "EHRSQL"
        }
      },
      {
        "value": "EHRSQL - # prompt tokens",
        "description": "EHRSQL is a benchmark designed to evaluate models on generating structured queries for clinical research. Each example includes a natural language question and a database schema, and the task is to produce an SQL query that would return the correct result for a biomedical research objective. This benchmark assesses a model's understanding of medical terminology, data structures, and query construction.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "EHRSQL"
        }
      },
      {
        "value": "EHRSQL - # output tokens",
        "description": "EHRSQL is a benchmark designed to evaluate models on generating structured queries for clinical research. Each example includes a natural language question and a database schema, and the task is to produce an SQL query that would return the correct result for a biomedical research objective. This benchmark assesses a model's understanding of medical terminology, data structures, and query construction.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "EHRSQL"
        }
      },
      {
        "value": "BMT-Status - # eval",
        "description": "BMT-Status is a benchmark composed of clinical notes and associated binary questions related to bone marrow transplant (BMT), hematopoietic stem cell transplant (HSCT), or hematopoietic cell transplant (HCT) status. The goal is to determine whether the patient received a subsequent transplant based on the provided clinical documentation.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "BMT-Status"
        }
      },
      {
        "value": "BMT-Status - # train",
        "description": "BMT-Status is a benchmark composed of clinical notes and associated binary questions related to bone marrow transplant (BMT), hematopoietic stem cell transplant (HSCT), or hematopoietic cell transplant (HCT) status. The goal is to determine whether the patient received a subsequent transplant based on the provided clinical documentation.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "BMT-Status"
        }
      },
      {
        "value": "BMT-Status - truncated",
        "description": "BMT-Status is a benchmark composed of clinical notes and associated binary questions related to bone marrow transplant (BMT), hematopoietic stem cell transplant (HSCT), or hematopoietic cell transplant (HCT) status. The goal is to determine whether the patient received a subsequent transplant based on the provided clinical documentation.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "BMT-Status"
        }
      },
      {
        "value": "BMT-Status - # prompt tokens",
        "description": "BMT-Status is a benchmark composed of clinical notes and associated binary questions related to bone marrow transplant (BMT), hematopoietic stem cell transplant (HSCT), or hematopoietic cell transplant (HCT) status. The goal is to determine whether the patient received a subsequent transplant based on the provided clinical documentation.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "BMT-Status"
        }
      },
      {
        "value": "BMT-Status - # output tokens",
        "description": "BMT-Status is a benchmark composed of clinical notes and associated binary questions related to bone marrow transplant (BMT), hematopoietic stem cell transplant (HSCT), or hematopoietic cell transplant (HCT) status. The goal is to determine whether the patient received a subsequent transplant based on the provided clinical documentation.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "BMT-Status"
        }
      },
      {
        "value": "RaceBias - # eval",
        "description": "RaceBias is a benchmark used to evaluate language models for racially biased or inappropriate content in medical question-answering scenarios. Each instance consists of a medical question and a model-generated response. The task is to classify whether the response contains race-based, harmful, or inaccurate content. This benchmark supports research into bias detection and fairness in clinical AI systems.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "RaceBias"
        }
      },
      {
        "value": "RaceBias - # train",
        "description": "RaceBias is a benchmark used to evaluate language models for racially biased or inappropriate content in medical question-answering scenarios. Each instance consists of a medical question and a model-generated response. The task is to classify whether the response contains race-based, harmful, or inaccurate content. This benchmark supports research into bias detection and fairness in clinical AI systems.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "RaceBias"
        }
      },
      {
        "value": "RaceBias - truncated",
        "description": "RaceBias is a benchmark used to evaluate language models for racially biased or inappropriate content in medical question-answering scenarios. Each instance consists of a medical question and a model-generated response. The task is to classify whether the response contains race-based, harmful, or inaccurate content. This benchmark supports research into bias detection and fairness in clinical AI systems.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "RaceBias"
        }
      },
      {
        "value": "RaceBias - # prompt tokens",
        "description": "RaceBias is a benchmark used to evaluate language models for racially biased or inappropriate content in medical question-answering scenarios. Each instance consists of a medical question and a model-generated response. The task is to classify whether the response contains race-based, harmful, or inaccurate content. This benchmark supports research into bias detection and fairness in clinical AI systems.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "RaceBias"
        }
      },
      {
        "value": "RaceBias - # output tokens",
        "description": "RaceBias is a benchmark used to evaluate language models for racially biased or inappropriate content in medical question-answering scenarios. Each instance consists of a medical question and a model-generated response. The task is to classify whether the response contains race-based, harmful, or inaccurate content. This benchmark supports research into bias detection and fairness in clinical AI systems.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "RaceBias"
        }
      },
      {
        "value": "N2C2-CT - # eval",
        "description": "A dataset that provides clinical notes and asks the model to classify whether the patient is a valid candidate for a provided clinical trial.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "N2C2-CT"
        }
      },
      {
        "value": "N2C2-CT - # train",
        "description": "A dataset that provides clinical notes and asks the model to classify whether the patient is a valid candidate for a provided clinical trial.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "N2C2-CT"
        }
      },
      {
        "value": "N2C2-CT - truncated",
        "description": "A dataset that provides clinical notes and asks the model to classify whether the patient is a valid candidate for a provided clinical trial.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "N2C2-CT"
        }
      },
      {
        "value": "N2C2-CT - # prompt tokens",
        "description": "A dataset that provides clinical notes and asks the model to classify whether the patient is a valid candidate for a provided clinical trial.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "N2C2-CT"
        }
      },
      {
        "value": "N2C2-CT - # output tokens",
        "description": "A dataset that provides clinical notes and asks the model to classify whether the patient is a valid candidate for a provided clinical trial.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "N2C2-CT"
        }
      },
      {
        "value": "MedHallu - # eval",
        "description": "MedHallu is a benchmark focused on evaluating factual correctness in biomedical question answering. Each instance contains a PubMed-derived knowledge snippet, a biomedical question, and a model-generated answer. The task is to classify whether the answer is factually correct or contains hallucinated (non-grounded) information. This benchmark is designed to assess the factual reliability of medical language models.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MedHallu"
        }
      },
      {
        "value": "MedHallu - # train",
        "description": "MedHallu is a benchmark focused on evaluating factual correctness in biomedical question answering. Each instance contains a PubMed-derived knowledge snippet, a biomedical question, and a model-generated answer. The task is to classify whether the answer is factually correct or contains hallucinated (non-grounded) information. This benchmark is designed to assess the factual reliability of medical language models.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MedHallu"
        }
      },
      {
        "value": "MedHallu - truncated",
        "description": "MedHallu is a benchmark focused on evaluating factual correctness in biomedical question answering. Each instance contains a PubMed-derived knowledge snippet, a biomedical question, and a model-generated answer. The task is to classify whether the answer is factually correct or contains hallucinated (non-grounded) information. This benchmark is designed to assess the factual reliability of medical language models.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MedHallu"
        }
      },
      {
        "value": "MedHallu - # prompt tokens",
        "description": "MedHallu is a benchmark focused on evaluating factual correctness in biomedical question answering. Each instance contains a PubMed-derived knowledge snippet, a biomedical question, and a model-generated answer. The task is to classify whether the answer is factually correct or contains hallucinated (non-grounded) information. This benchmark is designed to assess the factual reliability of medical language models.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MedHallu"
        }
      },
      {
        "value": "MedHallu - # output tokens",
        "description": "MedHallu is a benchmark focused on evaluating factual correctness in biomedical question answering. Each instance contains a PubMed-derived knowledge snippet, a biomedical question, and a model-generated answer. The task is to classify whether the answer is factually correct or contains hallucinated (non-grounded) information. This benchmark is designed to assess the factual reliability of medical language models.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MedHallu"
        }
      },
      {
        "value": "MIMIC-IV Billing Code - # eval",
        "description": "MIMIC-IV Billing Code is a benchmark derived from discharge summaries in the MIMIC-IV database, paired with their corresponding ICD-10 billing codes. The task requires models to extract structured billing codes based on free-text clinical notes, reflecting real-world hospital coding tasks for financial reimbursement.\n\n# eval: Number of evaluation instances.",
        "markdown": false,
        "metadata": {
          "metric": "# eval",
          "run_group": "MIMIC-IV Billing Code"
        }
      },
      {
        "value": "MIMIC-IV Billing Code - # train",
        "description": "MIMIC-IV Billing Code is a benchmark derived from discharge summaries in the MIMIC-IV database, paired with their corresponding ICD-10 billing codes. The task requires models to extract structured billing codes based on free-text clinical notes, reflecting real-world hospital coding tasks for financial reimbursement.\n\n# train: Number of training instances (e.g., in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "# train",
          "run_group": "MIMIC-IV Billing Code"
        }
      },
      {
        "value": "MIMIC-IV Billing Code - truncated",
        "description": "MIMIC-IV Billing Code is a benchmark derived from discharge summaries in the MIMIC-IV database, paired with their corresponding ICD-10 billing codes. The task requires models to extract structured billing codes based on free-text clinical notes, reflecting real-world hospital coding tasks for financial reimbursement.\n\ntruncated: Fraction of instances where the prompt itself was truncated (implies that there were no in-context examples).",
        "markdown": false,
        "metadata": {
          "metric": "truncated",
          "run_group": "MIMIC-IV Billing Code"
        }
      },
      {
        "value": "MIMIC-IV Billing Code - # prompt tokens",
        "description": "MIMIC-IV Billing Code is a benchmark derived from discharge summaries in the MIMIC-IV database, paired with their corresponding ICD-10 billing codes. The task requires models to extract structured billing codes based on free-text clinical notes, reflecting real-world hospital coding tasks for financial reimbursement.\n\n# prompt tokens: Number of tokens in the prompt.",
        "markdown": false,
        "metadata": {
          "metric": "# prompt tokens",
          "run_group": "MIMIC-IV Billing Code"
        }
      },
      {
        "value": "MIMIC-IV Billing Code - # output tokens",
        "description": "MIMIC-IV Billing Code is a benchmark derived from discharge summaries in the MIMIC-IV database, paired with their corresponding ICD-10 billing codes. The task requires models to extract structured billing codes based on free-text clinical notes, reflecting real-world hospital coding tasks for financial reimbursement.\n\n# output tokens: Actual number of output tokens.",
        "markdown": false,
        "metadata": {
          "metric": "# output tokens",
          "run_group": "MIMIC-IV Billing Code"
        }
      }
    ],
    "rows": [
      [
        {
          "value": "Claude 4.6 Opus",
          "description": "",
          "markdown": false
        },
        {
          "value": 1100.0,
          "description": "min=1100, mean=1100, max=1100, sum=1100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 713.3490909090909,
          "description": "min=713.349, mean=713.349, max=713.349, sum=713.349 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 59.67909090909091,
          "description": "min=59.679, mean=59.679, max=59.679, sum=59.679 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 789.3325526932084,
          "description": "min=789.333, mean=789.333, max=789.333, sum=789.333 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 989.0866510538642,
          "description": "min=989.087, mean=989.087, max=989.087, sum=989.087 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 311.64489112227807,
          "description": "min=311.645, mean=311.645, max=311.645, sum=311.645 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 208.17085427135677,
          "description": "min=208.171, mean=208.171, max=208.171, sum=208.171 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 444.0,
          "description": "min=396, mean=444, max=458, sum=2220 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=chemistry,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=medicine,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=pharmacology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=psychology,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=chemistry,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=medicine,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=pharmacology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=psychology,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=chemistry,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=medicine,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=pharmacology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=psychology,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 148.31642151470237,
          "description": "min=128.029, mean=148.316, max=193.538, sum=741.582 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=chemistry,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=medicine,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=pharmacology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=psychology,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 21.92201364780773,
          "description": "min=8.281, mean=21.922, max=55.581, sum=109.61 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=chemistry,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=medicine,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=pharmacology,model=anthropic_claude-opus-4.6",
            "head_qa:language=en,category=psychology,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 337.487012987013,
          "description": "min=337.487, mean=337.487, max=337.487, sum=337.487 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 169.90584415584416,
          "description": "min=169.906, mean=169.906, max=169.906, sum=169.906 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1273.0,
          "description": "min=1273, mean=1273, max=1273, sum=1273 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 259.46661429693637,
          "description": "min=259.467, mean=259.467, max=259.467, sum=259.467 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.9992144540455616,
          "description": "min=0.999, mean=0.999, max=0.999, sum=0.999 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4183.0,
          "description": "min=4183, mean=4183, max=4183, sum=4183 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 103.77384652163519,
          "description": "min=103.774, mean=103.774, max=103.774, sum=103.774 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1.0222328472388238,
          "description": "min=1.022, mean=1.022, max=1.022, sum=1.022 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1629.6166666666666,
          "description": "min=1629.617, mean=1629.617, max=1629.617, sum=1629.617 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 551.3666666666667,
          "description": "min=551.367, mean=551.367, max=551.367, sum=551.367 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 304.84375,
          "description": "min=304.844, mean=304.844, max=304.844, sum=304.844 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1008.5078125,
          "description": "min=1008.508, mean=1008.508, max=1008.508, sum=1008.508 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2977.0,
          "description": "min=2977, mean=2977, max=2977, sum=2977 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 219.65905273765534,
          "description": "min=219.659, mean=219.659, max=219.659, sum=219.659 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 88.4779979845482,
          "description": "min=88.478, mean=88.478, max=88.478, sum=88.478 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 644.34,
          "description": "min=644.34, mean=644.34, max=644.34, sum=644.34 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 185.48,
          "description": "min=185.48, mean=185.48, max=185.48, sum=185.48 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 22.712626995645863,
          "description": "min=22.713, mean=22.713, max=22.713, sum=22.713 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 330.79535558780844,
          "description": "min=330.795, mean=330.795, max=330.795, sum=330.795 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 316.1226,
          "description": "min=316.123, mean=316.123, max=316.123, sum=316.123 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1285.7478,
          "description": "min=1285.748, mean=1285.748, max=1285.748, sum=1285.748 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4054.0,
          "description": "min=3108, mean=4054, max=5000, sum=8108 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-opus-4.6",
            "med_dialog,subset=icliniq:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-opus-4.6",
            "med_dialog,subset=icliniq:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-opus-4.6",
            "med_dialog,subset=icliniq:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 261.02424646074644,
          "description": "min=242.377, mean=261.024, max=279.671, sum=522.048 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-opus-4.6",
            "med_dialog,subset=icliniq:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 64.76375276705276,
          "description": "min=63.784, mean=64.764, max=65.743, sum=129.528 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-opus-4.6",
            "med_dialog,subset=icliniq:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 157.864,
          "description": "min=157.864, mean=157.864, max=157.864, sum=157.864 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 38.16,
          "description": "min=38.16, mean=38.16, max=38.16, sum=38.16 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 488.74666666666667,
          "description": "min=488.747, mean=488.747, max=488.747, sum=488.747 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 163.2272,
          "description": "min=163.227, mean=163.227, max=163.227, sum=163.227 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 377.515,
          "description": "min=377.515, mean=377.515, max=377.515, sum=377.515 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2909.0,
          "description": "min=2909, mean=2909, max=2909, sum=2909 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 967.5768305259539,
          "description": "min=967.577, mean=967.577, max=967.577, sum=967.577 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 85.23616363011344,
          "description": "min=85.236, mean=85.236, max=85.236, sum=85.236 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 505.7098,
          "description": "min=505.71, mean=505.71, max=505.71, sum=505.71 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 473.0239520958084,
          "description": "min=473.024, mean=473.024, max=473.024, sum=473.024 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 4728.0,
          "description": "min=4728, mean=4728, max=4728, sum=14184 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 477.8830372250423,
          "description": "min=427.55, mean=477.883, max=554.55, sum=1433.649 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1.065707839819515,
          "description": "min=1, mean=1.066, max=1.197, sum=3.197 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-opus-4.6",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 2000.0,
          "description": "min=2000, mean=2000, max=2000, sum=2000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 699.848,
          "description": "min=699.848, mean=699.848, max=699.848, sum=699.848 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=anthropic_claude-opus-4.6"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Gemini 3.1 Pro (Preview)",
          "description": "",
          "markdown": false
        },
        {
          "value": 1100.0,
          "description": "min=1100, mean=1100, max=1100, sum=1100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 713.3490909090909,
          "description": "min=713.349, mean=713.349, max=713.349, sum=713.349 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 2.9145454545454546,
          "description": "min=2.915, mean=2.915, max=2.915, sum=2.915 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 789.3325526932084,
          "description": "min=789.333, mean=789.333, max=789.333, sum=789.333 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 532.1569086651053,
          "description": "min=532.157, mean=532.157, max=532.157, sum=532.157 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 311.64489112227807,
          "description": "min=311.645, mean=311.645, max=311.645, sum=311.645 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 19.42211055276382,
          "description": "min=19.422, mean=19.422, max=19.422, sum=19.422 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 444.0,
          "description": "min=396, mean=444, max=458, sum=2220 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=medicine,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=psychology,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=medicine,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=psychology,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=medicine,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=psychology,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 148.31642151470237,
          "description": "min=128.029, mean=148.316, max=193.538, sum=741.582 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=medicine,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=psychology,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4.112650868407721,
          "description": "min=1.959, mean=4.113, max=7.038, sum=20.563 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=medicine,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.1-pro-preview",
            "head_qa:language=en,category=psychology,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 337.487012987013,
          "description": "min=337.487, mean=337.487, max=337.487, sum=337.487 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 93.03896103896103,
          "description": "min=93.039, mean=93.039, max=93.039, sum=93.039 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 1273.0,
          "description": "min=1273, mean=1273, max=1273, sum=1273 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 259.46661429693637,
          "description": "min=259.467, mean=259.467, max=259.467, sum=259.467 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4183.0,
          "description": "min=4183, mean=4183, max=4183, sum=4183 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 103.77384652163519,
          "description": "min=103.774, mean=103.774, max=103.774, sum=103.774 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.9988046856323213,
          "description": "min=0.999, mean=0.999, max=0.999, sum=0.999 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 1677.2666666666667,
          "description": "min=1677.267, mean=1677.267, max=1677.267, sum=1677.267 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 568.925,
          "description": "min=568.925, mean=568.925, max=568.925, sum=568.925 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 294.734375,
          "description": "min=294.734, mean=294.734, max=294.734, sum=294.734 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 866.9375,
          "description": "min=866.938, mean=866.938, max=866.938, sum=866.938 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 2977.0,
          "description": "min=2977, mean=2977, max=2977, sum=2977 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 219.65905273765534,
          "description": "min=219.659, mean=219.659, max=219.659, sum=219.659 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 64.80349344978166,
          "description": "min=64.803, mean=64.803, max=64.803, sum=64.803 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 623.54,
          "description": "min=623.54, mean=623.54, max=623.54, sum=623.54 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 146.74,
          "description": "min=146.74, mean=146.74, max=146.74, sum=146.74 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 22.712626995645863,
          "description": "min=22.713, mean=22.713, max=22.713, sum=22.713 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 485.42525399129175,
          "description": "min=485.425, mean=485.425, max=485.425, sum=485.425 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 316.1226,
          "description": "min=316.123, mean=316.123, max=316.123, sum=316.123 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 863.8888,
          "description": "min=863.889, mean=863.889, max=863.889, sum=863.889 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4054.0,
          "description": "min=3108, mean=4054, max=5000, sum=8108 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.1-pro-preview",
            "med_dialog,subset=icliniq:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.1-pro-preview",
            "med_dialog,subset=icliniq:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.1-pro-preview",
            "med_dialog,subset=icliniq:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 261.02424646074644,
          "description": "min=242.377, mean=261.024, max=279.671, sum=522.048 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.1-pro-preview",
            "med_dialog,subset=icliniq:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 38.51701518661518,
          "description": "min=38.123, mean=38.517, max=38.911, sum=77.034 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.1-pro-preview",
            "med_dialog,subset=icliniq:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 157.864,
          "description": "min=157.864, mean=157.864, max=157.864, sum=157.864 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.9998,
          "description": "min=1.0, mean=1.0, max=1.0, sum=1.0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 38.16,
          "description": "min=38.16, mean=38.16, max=38.16, sum=38.16 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 719.4733333333334,
          "description": "min=719.473, mean=719.473, max=719.473, sum=719.473 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 163.2272,
          "description": "min=163.227, mean=163.227, max=163.227, sum=163.227 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 1.6456,
          "description": "min=1.646, mean=1.646, max=1.646, sum=1.646 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 377.515,
          "description": "min=377.515, mean=377.515, max=377.515, sum=377.515 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.999,
          "description": "min=0.999, mean=0.999, max=0.999, sum=0.999 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 2909.0,
          "description": "min=2909, mean=2909, max=2909, sum=2909 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 967.5768305259539,
          "description": "min=967.577, mean=967.577, max=967.577, sum=967.577 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 58.19147473358542,
          "description": "min=58.191, mean=58.191, max=58.191, sum=58.191 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 505.7098,
          "description": "min=505.71, mean=505.71, max=505.71, sum=505.71 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.9992,
          "description": "min=0.999, mean=0.999, max=0.999, sum=0.999 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 491.8862275449102,
          "description": "min=491.886, mean=491.886, max=491.886, sum=491.886 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 4728.0,
          "description": "min=4728, mean=4728, max=4728, sum=14184 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 477.8830372250423,
          "description": "min=427.55, mean=477.883, max=554.55, sum=1433.649 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.999224478285392,
          "description": "min=0.999, mean=0.999, max=1.0, sum=2.998 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.1-pro-preview",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 2000.0,
          "description": "min=2000, mean=2000, max=2000, sum=2000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 699.848,
          "description": "min=699.848, mean=699.848, max=699.848, sum=699.848 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=google_gemini-3.1-pro-preview"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Gemini-3.5-flash",
          "description": "",
          "markdown": false
        },
        {
          "value": 1100.0,
          "description": "min=1100, mean=1100, max=1100, sum=1100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 713.3490909090909,
          "description": "min=713.349, mean=713.349, max=713.349, sum=713.349 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2.7181818181818183,
          "description": "min=2.718, mean=2.718, max=2.718, sum=2.718 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 789.3325526932084,
          "description": "min=789.333, mean=789.333, max=789.333, sum=789.333 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 774.6229508196722,
          "description": "min=774.623, mean=774.623, max=774.623, sum=774.623 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 311.64489112227807,
          "description": "min=311.645, mean=311.645, max=311.645, sum=311.645 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 19.695142378559463,
          "description": "min=19.695, mean=19.695, max=19.695, sum=19.695 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 444.0,
          "description": "min=396, mean=444, max=458, sum=2220 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=medicine,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=psychology,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=medicine,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=psychology,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=medicine,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=psychology,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 148.31642151470237,
          "description": "min=128.029, mean=148.316, max=193.538, sum=741.582 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=medicine,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=psychology,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2.287403335918086,
          "description": "min=1.015, mean=2.287, max=3.808, sum=11.437 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=chemistry,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=medicine,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=pharmacology,model=google_gemini-3.5-flash",
            "head_qa:language=en,category=psychology,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 337.487012987013,
          "description": "min=337.487, mean=337.487, max=337.487, sum=337.487 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 36.188311688311686,
          "description": "min=36.188, mean=36.188, max=36.188, sum=36.188 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1273.0,
          "description": "min=1273, mean=1273, max=1273, sum=1273 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 259.46661429693637,
          "description": "min=259.467, mean=259.467, max=259.467, sum=259.467 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4183.0,
          "description": "min=4183, mean=4183, max=4183, sum=4183 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 103.77384652163519,
          "description": "min=103.774, mean=103.774, max=103.774, sum=103.774 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1.0325125508008606,
          "description": "min=1.033, mean=1.033, max=1.033, sum=1.033 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1629.6166666666666,
          "description": "min=1629.617, mean=1629.617, max=1629.617, sum=1629.617 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 633.6,
          "description": "min=633.6, mean=633.6, max=633.6, sum=633.6 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 304.84375,
          "description": "min=304.844, mean=304.844, max=304.844, sum=304.844 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 982.734375,
          "description": "min=982.734, mean=982.734, max=982.734, sum=982.734 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2977.0,
          "description": "min=2977, mean=2977, max=2977, sum=2977 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 219.65905273765534,
          "description": "min=219.659, mean=219.659, max=219.659, sum=219.659 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 77.80853207927444,
          "description": "min=77.809, mean=77.809, max=77.809, sum=77.809 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 644.34,
          "description": "min=644.34, mean=644.34, max=644.34, sum=644.34 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 191.74,
          "description": "min=191.74, mean=191.74, max=191.74, sum=191.74 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 22.712626995645863,
          "description": "min=22.713, mean=22.713, max=22.713, sum=22.713 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 557.4092888243831,
          "description": "min=557.409, mean=557.409, max=557.409, sum=557.409 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 316.1226,
          "description": "min=316.123, mean=316.123, max=316.123, sum=316.123 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 965.7082,
          "description": "min=965.708, mean=965.708, max=965.708, sum=965.708 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4054.0,
          "description": "min=3108, mean=4054, max=5000, sum=8108 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.5-flash",
            "med_dialog,subset=icliniq:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.5-flash",
            "med_dialog,subset=icliniq:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.5-flash",
            "med_dialog,subset=icliniq:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 261.02424646074644,
          "description": "min=242.377, mean=261.024, max=279.671, sum=522.048 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.5-flash",
            "med_dialog,subset=icliniq:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 40.61652792792793,
          "description": "min=40.356, mean=40.617, max=40.877, sum=81.233 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-3.5-flash",
            "med_dialog,subset=icliniq:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 157.864,
          "description": "min=157.864, mean=157.864, max=157.864, sum=157.864 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 38.16,
          "description": "min=38.16, mean=38.16, max=38.16, sum=38.16 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 774.68,
          "description": "min=774.68, mean=774.68, max=774.68, sum=774.68 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 163.2272,
          "description": "min=163.227, mean=163.227, max=163.227, sum=163.227 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 377.515,
          "description": "min=377.515, mean=377.515, max=377.515, sum=377.515 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2909.0,
          "description": "min=2909, mean=2909, max=2909, sum=2909 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 967.5768305259539,
          "description": "min=967.577, mean=967.577, max=967.577, sum=967.577 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 68.8102440701272,
          "description": "min=68.81, mean=68.81, max=68.81, sum=68.81 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 505.7098,
          "description": "min=505.71, mean=505.71, max=505.71, sum=505.71 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 473.0239520958084,
          "description": "min=473.024, mean=473.024, max=473.024, sum=473.024 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 4728.0,
          "description": "min=4728, mean=4728, max=4728, sum=14184 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 477.8830372250423,
          "description": "min=427.55, mean=477.883, max=554.55, sum=1433.649 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=3 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-3.5-flash",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 2000.0,
          "description": "min=2000, mean=2000, max=2000, sum=2000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 699.848,
          "description": "min=699.848, mean=699.848, max=699.848, sum=699.848 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 0.9995,
          "description": "min=1.0, mean=1.0, max=1.0, sum=1.0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-3.5-flash"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=google_gemini-3.5-flash"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "GPT-5.4 (2026-03-05)",
          "description": "",
          "markdown": false
        },
        {
          "value": 1100.0,
          "description": "min=1100, mean=1100, max=1100, sum=1100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 713.3490909090909,
          "description": "min=713.349, mean=713.349, max=713.349, sum=713.349 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 2.0163636363636366,
          "description": "min=2.016, mean=2.016, max=2.016, sum=2.016 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 789.3325526932084,
          "description": "min=789.333, mean=789.333, max=789.333, sum=789.333 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 631.4262295081967,
          "description": "min=631.426, mean=631.426, max=631.426, sum=631.426 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 311.64489112227807,
          "description": "min=311.645, mean=311.645, max=311.645, sum=311.645 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 17.088777219430487,
          "description": "min=17.089, mean=17.089, max=17.089, sum=17.089 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 444.0,
          "description": "min=396, mean=444, max=458, sum=2220 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 148.31642151470237,
          "description": "min=128.029, mean=148.316, max=193.538, sum=741.582 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.9921397379912664,
          "description": "min=0.961, mean=0.992, max=1, sum=4.961 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 337.487012987013,
          "description": "min=337.487, mean=337.487, max=337.487, sum=337.487 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 1273.0,
          "description": "min=1273, mean=1273, max=1273, sum=1273 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 259.46661429693637,
          "description": "min=259.467, mean=259.467, max=259.467, sum=259.467 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.9913589945011784,
          "description": "min=0.991, mean=0.991, max=0.991, sum=0.991 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4183.0,
          "description": "min=4183, mean=4183, max=4183, sum=4183 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 103.77384652163519,
          "description": "min=103.774, mean=103.774, max=103.774, sum=103.774 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.9569686827635668,
          "description": "min=0.957, mean=0.957, max=0.957, sum=0.957 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 1629.6166666666666,
          "description": "min=1629.617, mean=1629.617, max=1629.617, sum=1629.617 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 589.8333333333334,
          "description": "min=589.833, mean=589.833, max=589.833, sum=589.833 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 304.84375,
          "description": "min=304.844, mean=304.844, max=304.844, sum=304.844 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 713.0078125,
          "description": "min=713.008, mean=713.008, max=713.008, sum=713.008 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 2977.0,
          "description": "min=2977, mean=2977, max=2977, sum=2977 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 219.65905273765534,
          "description": "min=219.659, mean=219.659, max=219.659, sum=219.659 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 90.13604299630501,
          "description": "min=90.136, mean=90.136, max=90.136, sum=90.136 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 644.34,
          "description": "min=644.34, mean=644.34, max=644.34, sum=644.34 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 113.38,
          "description": "min=113.38, mean=113.38, max=113.38, sum=113.38 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 22.712626995645863,
          "description": "min=22.713, mean=22.713, max=22.713, sum=22.713 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 252.73730043541363,
          "description": "min=252.737, mean=252.737, max=252.737, sum=252.737 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 316.1226,
          "description": "min=316.123, mean=316.123, max=316.123, sum=316.123 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 815.6702,
          "description": "min=815.67, mean=815.67, max=815.67, sum=815.67 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4054.0,
          "description": "min=3108, mean=4054, max=5000, sum=8108 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 261.02424646074644,
          "description": "min=242.377, mean=261.024, max=279.671, sum=522.048 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 58.43843951093951,
          "description": "min=58.022, mean=58.438, max=58.855, sum=116.877 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 157.864,
          "description": "min=157.864, mean=157.864, max=157.864, sum=157.864 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 38.16,
          "description": "min=38.16, mean=38.16, max=38.16, sum=38.16 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 443.5933333333333,
          "description": "min=443.593, mean=443.593, max=443.593, sum=443.593 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 163.2272,
          "description": "min=163.227, mean=163.227, max=163.227, sum=163.227 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.9962,
          "description": "min=0.996, mean=0.996, max=0.996, sum=0.996 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 377.515,
          "description": "min=377.515, mean=377.515, max=377.515, sum=377.515 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.98,
          "description": "min=0.98, mean=0.98, max=0.98, sum=0.98 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 2909.0,
          "description": "min=2909, mean=2909, max=2909, sum=2909 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 967.5768305259539,
          "description": "min=967.577, mean=967.577, max=967.577, sum=967.577 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 94.47507734616707,
          "description": "min=94.475, mean=94.475, max=94.475, sum=94.475 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 505.7098,
          "description": "min=505.71, mean=505.71, max=505.71, sum=505.71 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 473.0239520958084,
          "description": "min=473.024, mean=473.024, max=473.024, sum=473.024 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 4728.0,
          "description": "min=4728, mean=4728, max=4728, sum=14184 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 477.8830372250423,
          "description": "min=427.55, mean=477.883, max=554.55, sum=1433.649 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.9917512690355329,
          "description": "min=0.98, mean=0.992, max=1, sum=2.975 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4"
          ]
        },
        {
          "value": 2000.0,
          "description": "min=2000, mean=2000, max=2000, sum=2000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 699.848,
          "description": "min=699.848, mean=699.848, max=699.848, sum=699.848 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 0.821,
          "description": "min=0.821, mean=0.821, max=0.821, sum=0.821 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=openai_gpt-5.4"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "GPT-5.4 mini",
          "description": "",
          "markdown": false
        },
        {
          "value": 1100.0,
          "description": "min=1100, mean=1100, max=1100, sum=1100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 713.3490909090909,
          "description": "min=713.349, mean=713.349, max=713.349, sum=713.349 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 2.1063636363636364,
          "description": "min=2.106, mean=2.106, max=2.106, sum=2.106 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 789.3325526932084,
          "description": "min=789.333, mean=789.333, max=789.333, sum=789.333 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 366.0327868852459,
          "description": "min=366.033, mean=366.033, max=366.033, sum=366.033 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 311.64489112227807,
          "description": "min=311.645, mean=311.645, max=311.645, sum=311.645 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 15.618090452261306,
          "description": "min=15.618, mean=15.618, max=15.618, sum=15.618 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 444.0,
          "description": "min=396, mean=444, max=458, sum=2220 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 148.31642151470237,
          "description": "min=128.029, mean=148.316, max=193.538, sum=741.582 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=5 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=chemistry,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=medicine,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=pharmacology,model=openai_gpt-5.4-mini",
            "head_qa:language=en,category=psychology,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 337.487012987013,
          "description": "min=337.487, mean=337.487, max=337.487, sum=337.487 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1273.0,
          "description": "min=1273, mean=1273, max=1273, sum=1273 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 259.46661429693637,
          "description": "min=259.467, mean=259.467, max=259.467, sum=259.467 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4183.0,
          "description": "min=4183, mean=4183, max=4183, sum=4183 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 103.77384652163519,
          "description": "min=103.774, mean=103.774, max=103.774, sum=103.774 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.9997609371264643,
          "description": "min=1.0, mean=1.0, max=1.0, sum=1.0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1629.6166666666666,
          "description": "min=1629.617, mean=1629.617, max=1629.617, sum=1629.617 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 525.4833333333333,
          "description": "min=525.483, mean=525.483, max=525.483, sum=525.483 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 304.84375,
          "description": "min=304.844, mean=304.844, max=304.844, sum=304.844 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 459.3359375,
          "description": "min=459.336, mean=459.336, max=459.336, sum=459.336 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 2977.0,
          "description": "min=2977, mean=2977, max=2977, sum=2977 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 219.65905273765534,
          "description": "min=219.659, mean=219.659, max=219.659, sum=219.659 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 93.68626133691636,
          "description": "min=93.686, mean=93.686, max=93.686, sum=93.686 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 644.34,
          "description": "min=644.34, mean=644.34, max=644.34, sum=644.34 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 113.06,
          "description": "min=113.06, mean=113.06, max=113.06, sum=113.06 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 22.712626995645863,
          "description": "min=22.713, mean=22.713, max=22.713, sum=22.713 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 187.1465892597968,
          "description": "min=187.147, mean=187.147, max=187.147, sum=187.147 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 316.1226,
          "description": "min=316.123, mean=316.123, max=316.123, sum=316.123 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 723.9308,
          "description": "min=723.931, mean=723.931, max=723.931, sum=723.931 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4054.0,
          "description": "min=3108, mean=4054, max=5000, sum=8108 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4-mini",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4-mini",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4-mini",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 261.02424646074644,
          "description": "min=242.377, mean=261.024, max=279.671, sum=522.048 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4-mini",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 55.69369189189189,
          "description": "min=55.284, mean=55.694, max=56.104, sum=111.387 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=openai_gpt-5.4-mini",
            "med_dialog,subset=icliniq:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 157.864,
          "description": "min=157.864, mean=157.864, max=157.864, sum=157.864 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 38.16,
          "description": "min=38.16, mean=38.16, max=38.16, sum=38.16 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 348.52,
          "description": "min=348.52, mean=348.52, max=348.52, sum=348.52 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 163.2272,
          "description": "min=163.227, mean=163.227, max=163.227, sum=163.227 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.9998,
          "description": "min=1.0, mean=1.0, max=1.0, sum=1.0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 377.515,
          "description": "min=377.515, mean=377.515, max=377.515, sum=377.515 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 2909.0,
          "description": "min=2909, mean=2909, max=2909, sum=2909 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 967.5768305259539,
          "description": "min=967.577, mean=967.577, max=967.577, sum=967.577 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 76.95840495015469,
          "description": "min=76.958, mean=76.958, max=76.958, sum=76.958 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 505.7098,
          "description": "min=505.71, mean=505.71, max=505.71, sum=505.71 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med,stop=none:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 473.0239520958084,
          "description": "min=473.024, mean=473.024, max=473.024, sum=473.024 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 4728.0,
          "description": "min=4728, mean=4728, max=4728, sum=14184 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 477.8830372250423,
          "description": "min=427.55, mean=477.883, max=554.55, sum=1433.649 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.9998589960518895,
          "description": "min=1.0, mean=1.0, max=1, sum=3.0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=openai_gpt-5.4-mini",
            "n2c2_ct_matching:subject=CREATININE,model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 2000.0,
          "description": "min=2000, mean=2000, max=2000, sum=2000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 699.848,
          "description": "min=699.848, mean=699.848, max=699.848, sum=699.848 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=openai_gpt-5.4-mini"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Claude 3.7 Sonnet (20250219)",
          "description": "",
          "markdown": false
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 579.799,
          "description": "min=579.799, mean=579.799, max=579.799, sum=579.799 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 2.091,
          "description": "min=2.091, mean=2.091, max=2.091, sum=2.091 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 794.0585480093677,
          "description": "min=794.059, mean=794.059, max=794.059, sum=794.059 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 320.5175879396985,
          "description": "min=320.518, mean=320.518, max=320.518, sum=320.518 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 17.23785594639866,
          "description": "min=17.238, mean=17.238, max=17.238, sum=17.238 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 163.022,
          "description": "min=163.022, mean=163.022, max=163.022, sum=163.022 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 345.6266233766234,
          "description": "min=345.627, mean=345.627, max=345.627, sum=345.627 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 503.0,
          "description": "min=503, mean=503, max=503, sum=503 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 264.4373757455268,
          "description": "min=264.437, mean=264.437, max=264.437, sum=264.437 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 107.214,
          "description": "min=107.214, mean=107.214, max=107.214, sum=107.214 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1674.0583333333334,
          "description": "min=1674.058, mean=1674.058, max=1674.058, sum=1674.058 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 309.640625,
          "description": "min=309.641, mean=309.641, max=309.641, sum=309.641 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 212.638,
          "description": "min=212.638, mean=212.638, max=212.638, sum=212.638 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 653.94,
          "description": "min=653.94, mean=653.94, max=653.94, sum=653.94 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 25.08708272859216,
          "description": "min=25.087, mean=25.087, max=25.087, sum=25.087 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 361.0,
          "description": "min=361, mean=361, max=361, sum=361 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 2571.130193905817,
          "description": "min=2571.13, mean=2571.13, max=2571.13, sum=2571.13 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=2000 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "med_dialog,subset=icliniq:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "med_dialog,subset=icliniq:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "med_dialog,subset=icliniq:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 267.246,
          "description": "min=244.906, mean=267.246, max=289.586, sum=534.492 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "med_dialog,subset=icliniq:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "med_dialog,subset=icliniq:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 2062.048,
          "description": "min=2062.048, mean=2062.048, max=2062.048, sum=2062.048 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 40.62,
          "description": "min=40.62, mean=40.62, max=40.62, sum=40.62 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 300.0,
          "description": "min=300, mean=300, max=300, sum=300 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 188.42333333333335,
          "description": "min=188.423, mean=188.423, max=188.423, sum=188.423 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 375.302,
          "description": "min=375.302, mean=375.302, max=375.302, sum=375.302 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1126.031,
          "description": "min=1126.031, mean=1126.031, max=1126.031, sum=1126.031 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 81.993,
          "description": "min=81.993, mean=81.993, max=81.993, sum=81.993 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 220.0,
          "description": "min=220, mean=220, max=220, sum=220 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 2041.9727272727273,
          "description": "min=2041.973, mean=2041.973, max=2041.973, sum=2041.973 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 490.29940119760477,
          "description": "min=490.299, mean=490.299, max=490.299, sum=490.299 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 86.0,
          "description": "min=86, mean=86, max=86, sum=258 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 5753.391472868217,
          "description": "min=5702.058, mean=5753.391, max=5830.058, sum=17260.174 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219",
            "n2c2_ct_matching:subject=CREATININE,model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 715.166,
          "description": "min=715.166, mean=715.166, max=715.166, sum=715.166 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=anthropic_claude-3-7-sonnet-20250219,model_deployment=stanfordhealthcare_claude-3-7-sonnet-20250219"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "DeepSeek R1",
          "description": "",
          "markdown": false
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 551.78,
          "description": "min=551.78, mean=551.78, max=551.78, sum=551.78 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 756.0468384074942,
          "description": "min=756.047, mean=756.047, max=756.047, sum=756.047 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 596.5573770491803,
          "description": "min=596.557, mean=596.557, max=596.557, sum=596.557 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 304.89447236180905,
          "description": "min=304.894, mean=304.894, max=304.894, sum=304.894 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 146.889,
          "description": "min=146.889, mean=146.889, max=146.889, sum=146.889 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 330.1655844155844,
          "description": "min=330.166, mean=330.166, max=330.166, sum=330.166 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 503.0,
          "description": "min=503, mean=503, max=503, sum=503 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 251.93041749502981,
          "description": "min=251.93, mean=251.93, max=251.93, sum=251.93 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.9582504970178927,
          "description": "min=0.958, mean=0.958, max=0.958, sum=0.958 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 100.817,
          "description": "min=100.817, mean=100.817, max=100.817, sum=100.817 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.968,
          "description": "min=0.968, mean=0.968, max=0.968, sum=0.968 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1613.1666666666667,
          "description": "min=1613.167, mean=1613.167, max=1613.167, sum=1613.167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 465.925,
          "description": "min=465.925, mean=465.925, max=465.925, sum=465.925 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 289.765625,
          "description": "min=289.766, mean=289.766, max=289.766, sum=289.766 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 590.8515625,
          "description": "min=590.852, mean=590.852, max=590.852, sum=590.852 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 199.291,
          "description": "min=199.291, mean=199.291, max=199.291, sum=199.291 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 94.074,
          "description": "min=94.074, mean=94.074, max=94.074, sum=94.074 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 635.72,
          "description": "min=635.72, mean=635.72, max=635.72, sum=635.72 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 149.09,
          "description": "min=149.09, mean=149.09, max=149.09, sum=149.09 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 22.378809869375907,
          "description": "min=22.379, mean=22.379, max=22.379, sum=22.379 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 285.86066763425254,
          "description": "min=285.861, mean=285.861, max=285.861, sum=285.861 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 361.0,
          "description": "min=361, mean=361, max=361, sum=361 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 2541.786703601108,
          "description": "min=2541.787, mean=2541.787, max=2541.787, sum=2541.787 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 699.0747922437673,
          "description": "min=699.075, mean=699.075, max=699.075, sum=699.075 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=2000 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 260.207,
          "description": "min=239.027, mean=260.207, max=281.387, sum=520.414 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 58.4075,
          "description": "min=57.567, mean=58.407, max=59.248, sum=116.815 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 2017.948,
          "description": "min=2017.948, mean=2017.948, max=2017.948, sum=2017.948 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 37.50666666666667,
          "description": "min=37.507, mean=37.507, max=37.507, sum=37.507 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 421.92,
          "description": "min=421.92, mean=421.92, max=421.92, sum=421.92 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 300.0,
          "description": "min=300, mean=300, max=300, sum=300 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 183.22666666666666,
          "description": "min=183.227, mean=183.227, max=183.227, sum=183.227 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 364.658,
          "description": "min=364.658, mean=364.658, max=364.658, sum=364.658 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1002.984,
          "description": "min=1002.984, mean=1002.984, max=1002.984, sum=1002.984 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 220.0,
          "description": "min=220, mean=220, max=220, sum=220 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 2042.3363636363636,
          "description": "min=2042.336, mean=2042.336, max=2042.336, sum=2042.336 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1.0363636363636364,
          "description": "min=1.036, mean=1.036, max=1.036, sum=1.036 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 466.94610778443115,
          "description": "min=466.946, mean=466.946, max=466.946, sum=466.946 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 86.0,
          "description": "min=86, mean=86, max=86, sum=258 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 5437.860465116279,
          "description": "min=5387.86, mean=5437.86, max=5515.86, sum=16313.581 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.9922480620155039,
          "description": "min=0.977, mean=0.992, max=1, sum=2.977 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 694.905,
          "description": "min=694.905, mean=694.905, max=694.905, sum=694.905 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:num_output_tokens=4000,model=deepseek-ai_deepseek-r1,model_deployment=stanfordhealthcare_deepseek-r1"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Gemini 2.0 Flash",
          "description": "",
          "markdown": false
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 586.742,
          "description": "min=586.742, mean=586.742, max=586.742, sum=586.742 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 759.2295081967213,
          "description": "min=759.23, mean=759.23, max=759.23, sum=759.23 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 327.09212730318256,
          "description": "min=327.092, mean=327.092, max=327.092, sum=327.092 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 153.899,
          "description": "min=153.899, mean=153.899, max=153.899, sum=153.899 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 347.45454545454544,
          "description": "min=347.455, mean=347.455, max=347.455, sum=347.455 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 503.0,
          "description": "min=503, mean=503, max=503, sum=503 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 263.48111332007954,
          "description": "min=263.481, mean=263.481, max=263.481, sum=263.481 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 102.476,
          "description": "min=102.476, mean=102.476, max=102.476, sum=102.476 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1677.5333333333333,
          "description": "min=1677.533, mean=1677.533, max=1677.533, sum=1677.533 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 289.015625,
          "description": "min=289.016, mean=289.016, max=289.016, sum=289.016 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 197.6,
          "description": "min=197.6, mean=197.6, max=197.6, sum=197.6 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 615.41,
          "description": "min=615.41, mean=615.41, max=615.41, sum=615.41 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 23.423802612481857,
          "description": "min=23.424, mean=23.424, max=23.424, sum=23.424 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 361.0,
          "description": "min=361, mean=361, max=361, sum=361 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2606.409972299169,
          "description": "min=2606.41, mean=2606.41, max=2606.41, sum=2606.41 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=2000 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "med_dialog,subset=icliniq:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "med_dialog,subset=icliniq:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "med_dialog,subset=icliniq:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 261.433,
          "description": "min=238.927, mean=261.433, max=283.939, sum=522.866 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "med_dialog,subset=icliniq:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "med_dialog,subset=icliniq:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2089.742,
          "description": "min=2089.742, mean=2089.742, max=2089.742, sum=2089.742 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 38.906666666666666,
          "description": "min=38.907, mean=38.907, max=38.907, sum=38.907 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 300.0,
          "description": "min=300, mean=300, max=300, sum=300 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 187.90666666666667,
          "description": "min=187.907, mean=187.907, max=187.907, sum=187.907 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 383.966,
          "description": "min=383.966, mean=383.966, max=383.966, sum=383.966 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1112.729,
          "description": "min=1112.729, mean=1112.729, max=1112.729, sum=1112.729 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 220.0,
          "description": "min=220, mean=220, max=220, sum=220 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 2157.181818181818,
          "description": "min=2157.182, mean=2157.182, max=2157.182, sum=2157.182 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 490.7544910179641,
          "description": "min=490.754, mean=490.754, max=490.754, sum=490.754 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 86.0,
          "description": "min=86, mean=86, max=86, sum=258 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 5833.290697674419,
          "description": "min=5783.291, mean=5833.291, max=5908.291, sum=17499.872 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001",
            "n2c2_ct_matching:subject=CREATININE,model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 718.045,
          "description": "min=718.045, mean=718.045, max=718.045, sum=718.045 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=google_gemini-2.0-flash-001,model_deployment=stanfordhealthcare_gemini-2.0-flash-001"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Gemini 2.5 Pro (05-06 preview)",
          "description": "",
          "markdown": false
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 586.742,
          "description": "min=586.742, mean=586.742, max=586.742, sum=586.742 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 759.2295081967213,
          "description": "min=759.23, mean=759.23, max=759.23, sum=759.23 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 327.09212730318256,
          "description": "min=327.092, mean=327.092, max=327.092, sum=327.092 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 153.899,
          "description": "min=153.899, mean=153.899, max=153.899, sum=153.899 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 347.45454545454544,
          "description": "min=347.455, mean=347.455, max=347.455, sum=347.455 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 503.0,
          "description": "min=503, mean=503, max=503, sum=503 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 263.48111332007954,
          "description": "min=263.481, mean=263.481, max=263.481, sum=263.481 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 102.476,
          "description": "min=102.476, mean=102.476, max=102.476, sum=102.476 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1677.5333333333333,
          "description": "min=1677.533, mean=1677.533, max=1677.533, sum=1677.533 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 289.015625,
          "description": "min=289.016, mean=289.016, max=289.016, sum=289.016 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 197.6,
          "description": "min=197.6, mean=197.6, max=197.6, sum=197.6 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 615.41,
          "description": "min=615.41, mean=615.41, max=615.41, sum=615.41 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 23.423802612481857,
          "description": "min=23.424, mean=23.424, max=23.424, sum=23.424 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 361.0,
          "description": "min=361, mean=361, max=361, sum=361 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 2606.409972299169,
          "description": "min=2606.41, mean=2606.41, max=2606.41, sum=2606.41 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=2000 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 261.433,
          "description": "min=238.927, mean=261.433, max=283.939, sum=522.866 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "med_dialog,subset=icliniq:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 2089.742,
          "description": "min=2089.742, mean=2089.742, max=2089.742, sum=2089.742 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 38.906666666666666,
          "description": "min=38.907, mean=38.907, max=38.907, sum=38.907 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 300.0,
          "description": "min=300, mean=300, max=300, sum=300 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 177.90666666666667,
          "description": "min=177.907, mean=177.907, max=177.907, sum=177.907 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 500.0,
          "description": "min=500, mean=500, max=500, sum=500 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 388.094,
          "description": "min=388.094, mean=388.094, max=388.094, sum=388.094 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1112.729,
          "description": "min=1112.729, mean=1112.729, max=1112.729, sum=1112.729 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 220.0,
          "description": "min=220, mean=220, max=220, sum=220 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 2157.181818181818,
          "description": "min=2157.182, mean=2157.182, max=2157.182, sum=2157.182 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 490.7544910179641,
          "description": "min=490.754, mean=490.754, max=490.754, sum=490.754 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 86.0,
          "description": "min=86, mean=86, max=86, sum=258 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 5833.290697674419,
          "description": "min=5783.291, mean=5833.291, max=5908.291, sum=17499.872 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=ADVANCED-CAD,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06",
            "n2c2_ct_matching:subject=CREATININE,num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 705.045,
          "description": "min=705.045, mean=705.045, max=705.045, sum=705.045 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:num_output_tokens=4000,model=google_gemini-2.5-pro-preview-05-06,model_deployment=stanfordhealthcare_gemini-2.5-pro-preview-05-06"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Llama 3.3 Instruct (70B)",
          "description": "",
          "markdown": false
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 584.571,
          "description": "min=584.571, mean=584.571, max=584.571, sum=584.571 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 4.27,
          "description": "min=4.27, mean=4.27, max=4.27, sum=4.27 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 788.0538641686182,
          "description": "min=788.054, mean=788.054, max=788.054, sum=788.054 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 443.0234192037471,
          "description": "min=443.023, mean=443.023, max=443.023, sum=443.023 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 311.5812395309883,
          "description": "min=311.581, mean=311.581, max=311.581, sum=311.581 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 171.64489112227807,
          "description": "min=171.645, mean=171.645, max=171.645, sum=171.645 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 150.567,
          "description": "min=150.567, mean=150.567, max=150.567, sum=150.567 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=None,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 337.39935064935065,
          "description": "min=337.399, mean=337.399, max=337.399, sum=337.399 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 503.0,
          "description": "min=503, mean=503, max=503, sum=503 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 260.27435387673955,
          "description": "min=260.274, mean=260.274, max=260.274, sum=260.274 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.9960238568588469,
          "description": "min=0.996, mean=0.996, max=0.996, sum=0.996 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 103.403,
          "description": "min=103.403, mean=103.403, max=103.403, sum=103.403 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.998,
          "description": "min=0.998, mean=0.998, max=0.998, sum=0.998 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1629.5833333333333,
          "description": "min=1629.583, mean=1629.583, max=1629.583, sum=1629.583 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 428.6666666666667,
          "description": "min=428.667, mean=428.667, max=428.667, sum=428.667 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 304.1640625,
          "description": "min=304.164, mean=304.164, max=304.164, sum=304.164 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 456.828125,
          "description": "min=456.828, mean=456.828, max=456.828, sum=456.828 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 210.318,
          "description": "min=210.318, mean=210.318, max=210.318, sum=210.318 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 54.236,
          "description": "min=54.236, mean=54.236, max=54.236, sum=54.236 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 643.16,
          "description": "min=643.16, mean=643.16, max=643.16, sum=643.16 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 118.91,
          "description": "min=118.91, mean=118.91, max=118.91, sum=118.91 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 22.6966618287373,
          "description": "min=22.697, mean=22.697, max=22.697, sum=22.697 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 295.51959361393324,
          "description": "min=295.52, mean=295.52, max=295.52, sum=295.52 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 361.0,
          "description": "min=361, mean=361, max=361, sum=361 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 2616.9639889196674,
          "description": "min=2616.964, mean=2616.964, max=2616.964, sum=2616.964 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 255.98060941828254,
          "description": "min=255.981, mean=255.981, max=255.981, sum=255.981 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=2000 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "med_dialog,subset=icliniq:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "med_dialog,subset=icliniq:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "med_dialog,subset=icliniq:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 261.095,
          "description": "min=239.772, mean=261.095, max=282.418, sum=522.19 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "med_dialog,subset=icliniq:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 51.19,
          "description": "min=50.736, mean=51.19, max=51.644, sum=102.38 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "med_dialog,subset=icliniq:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 2070.752,
          "description": "min=2070.752, mean=2070.752, max=2070.752, sum=2070.752 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 38.16,
          "description": "min=38.16, mean=38.16, max=38.16, sum=38.16 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 486.67333333333335,
          "description": "min=486.673, mean=486.673, max=486.673, sum=486.673 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 300.0,
          "description": "min=300, mean=300, max=300, sum=300 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 183.76,
          "description": "min=183.76, mean=183.76, max=183.76, sum=183.76 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 376.774,
          "description": "min=376.774, mean=376.774, max=376.774, sum=376.774 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.102,
          "description": "min=1.102, mean=1.102, max=1.102, sum=1.102 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 962.504,
          "description": "min=962.504, mean=962.504, max=962.504, sum=962.504 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 61.127,
          "description": "min=61.127, mean=61.127, max=61.127, sum=61.127 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 220.0,
          "description": "min=220, mean=220, max=220, sum=220 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 2091.018181818182,
          "description": "min=2091.018, mean=2091.018, max=2091.018, sum=2091.018 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 472.91616766467064,
          "description": "min=472.916, mean=472.916, max=472.916, sum=472.916 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 86.0,
          "description": "min=86, mean=86, max=86, sum=258 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=CREATININE,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=CREATININE,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=CREATININE,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 5588.84496124031,
          "description": "min=5538.512, mean=5588.845, max=5665.512, sum=16766.535 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=CREATININE,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=3 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct",
            "n2c2_ct_matching:subject=CREATININE,model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 708.639,
          "description": "min=708.639, mean=708.639, max=708.639, sum=708.639 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=meta_llama-3.3-70b-instruct,model_deployment=stanfordhealthcare_llama-3.3-70b-instruct"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ],
      [
        {
          "value": "Muse Spark (2026-04-08)",
          "description": "",
          "markdown": false
        },
        {
          "value": 1100.0,
          "description": "min=1100, mean=1100, max=1100, sum=1100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 686.8127272727273,
          "description": "min=686.813, mean=686.813, max=686.813, sum=686.813 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medcalc_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 427.0,
          "description": "min=427, mean=427, max=427, sum=427 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 761.7330210772834,
          "description": "min=761.733, mean=761.733, max=761.733, sum=761.733 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1242.2833723653396,
          "description": "min=1242.283, mean=1242.283, max=1242.283, sum=1242.283 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_replicate:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 597.0,
          "description": "min=597, mean=597, max=597, sum=597 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 306.8542713567839,
          "description": "min=306.854, mean=306.854, max=306.854, sum=306.854 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 17.743718592964825,
          "description": "min=17.744, mean=17.744, max=17.744, sum=17.744 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medec:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 444.0,
          "description": "min=396, mean=444, max=458, sum=2220 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=chemistry,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=medicine,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=pharmacology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=psychology,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=chemistry,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=medicine,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=pharmacology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=psychology,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=chemistry,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=medicine,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=pharmacology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=psychology,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 145.37351267405694,
          "description": "min=125.405, mean=145.374, max=188.977, sum=726.868 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=chemistry,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=medicine,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=pharmacology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=psychology,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (5)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "head_qa:language=en,category=biology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=chemistry,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=medicine,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=pharmacology,model=meta_quiet_sand_v1alpha",
            "head_qa:language=en,category=psychology,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 308.0,
          "description": "min=308, mean=308, max=308, sum=308 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 331.59415584415586,
          "description": "min=331.594, mean=331.594, max=331.594, sum=331.594 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1.0292207792207793,
          "description": "min=1.029, mean=1.029, max=1.029, sum=1.029 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medbullets:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1273.0,
          "description": "min=1273, mean=1273, max=1273, sum=1273 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 254.17910447761193,
          "description": "min=254.179, mean=254.179, max=254.179, sum=254.179 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4183.0,
          "description": "min=4183, mean=4183, max=4183, sum=4183 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 101.77599808749702,
          "description": "min=101.776, mean=101.776, max=101.776, sum=101.776 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_mcqa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 120.0,
          "description": "min=120, mean=120, max=120, sum=120 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1574.3166666666666,
          "description": "min=1574.317, mean=1574.317, max=1574.317, sum=1574.317 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 720.075,
          "description": "min=720.075, mean=720.075, max=720.075, sum=720.075 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "aci_bench:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 128.0,
          "description": "min=128, mean=128, max=128, sum=128 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 292.109375,
          "description": "min=292.109, mean=292.109, max=292.109, sum=292.109 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1214.9453125,
          "description": "min=1214.945, mean=1214.945, max=1214.945, sum=1214.945 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mtsamples_procedures:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 2977.0,
          "description": "min=2977, mean=2977, max=2977, sum=2977 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 212.5979173664763,
          "description": "min=212.598, mean=212.598, max=212.598, sum=212.598 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 90.84145112529391,
          "description": "min=90.841, mean=90.841, max=90.841, sum=90.841 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_rrs:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 100.0,
          "description": "min=100, mean=100, max=100, sum=100 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 622.39,
          "description": "min=622.39, mean=622.39, max=622.39, sum=622.39 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 367.99,
          "description": "min=367.99, mean=367.99, max=367.99, sum=367.99 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimic_bhc:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 689.0,
          "description": "min=689, mean=689, max=689, sum=689 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 22.420899854862117,
          "description": "min=22.421, mean=22.421, max=22.421, sum=22.421 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medication_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 305.1294,
          "description": "min=305.129, mean=305.129, max=305.129, sum=305.129 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "starr_patient_instructions:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4054.0,
          "description": "min=3108, mean=4054, max=5000, sum=8108 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_quiet_sand_v1alpha",
            "med_dialog,subset=icliniq:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_quiet_sand_v1alpha",
            "med_dialog,subset=icliniq:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_quiet_sand_v1alpha",
            "med_dialog,subset=icliniq:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 259.65073204633205,
          "description": "min=240.209, mean=259.651, max=279.093, sum=519.301 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_quiet_sand_v1alpha",
            "med_dialog,subset=icliniq:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "med_dialog,subset=healthcaremagic:model=meta_quiet_sand_v1alpha",
            "med_dialog,subset=icliniq:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 157.7644,
          "description": "min=157.764, mean=157.764, max=157.764, sum=157.764 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_conf_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 150.0,
          "description": "min=150, mean=150, max=150, sum=150 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 37.63333333333333,
          "description": "min=37.633, mean=37.633, max=37.633, sum=37.633 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 801.4533333333334,
          "description": "min=801.453, mean=801.453, max=801.453, sum=801.453 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medi_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 161.0212,
          "description": "min=161.021, mean=161.021, max=161.021, sum=161.021 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_proxy_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1000.0,
          "description": "min=1000, mean=1000, max=1000, sum=1000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 367.191,
          "description": "min=367.191, mean=367.191, max=367.191, sum=367.191 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "pubmed_qa:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 2909.0,
          "description": "min=2909, mean=2909, max=2909, sum=5818 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 957.1285665177037,
          "description": "min=955.129, mean=957.129, max=959.129, sum=1914.257 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 34.05345479546236,
          "description": "min=0, mean=34.053, max=68.107, sum=68.107 (2)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "ehr_sql:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 490.1558,
          "description": "min=490.156, mean=490.156, max=490.156, sum=490.156 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "shc_bmt_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 167.0,
          "description": "min=167, mean=167, max=167, sum=167 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 466.80838323353294,
          "description": "min=466.808, mean=466.808, max=466.808, sum=466.808 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "race_based_med:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 4728.0,
          "description": "min=4728, mean=4728, max=4728, sum=14184 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=CREATININE,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=CREATININE,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=CREATININE,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 464.1546108291032,
          "description": "min=415.155, mean=464.155, max=540.155, sum=1392.464 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=CREATININE,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (3)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "n2c2_ct_matching:subject=ABDOMINAL,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=ADVANCED-CAD,model=meta_quiet_sand_v1alpha",
            "n2c2_ct_matching:subject=CREATININE,model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 2000.0,
          "description": "min=2000, mean=2000, max=2000, sum=2000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 0.0,
          "description": "min=0, mean=0, max=0, sum=0 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 693.254,
          "description": "min=693.254, mean=693.254, max=693.254, sum=693.254 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 1.0,
          "description": "min=1, mean=1, max=1, sum=1 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "medhallu:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "value": 5000.0,
          "description": "min=5000, mean=5000, max=5000, sum=5000 (1)",
          "style": {},
          "markdown": false,
          "run_spec_names": [
            "mimiciv_billing_code:model=meta_quiet_sand_v1alpha"
          ]
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        },
        {
          "description": "1 matching runs, but no matching metrics",
          "markdown": false
        }
      ]
    ],
    "links": [
      {
        "text": "LaTeX",
        "href": "benchmarks/releases/v5.0.0/groups/latex/medhelm_scenarios_general_information.tex"
      },
      {
        "text": "JSON",
        "href": "benchmarks/releases/v5.0.0/groups/json/medhelm_scenarios_general_information.json"
      }
    ],
    "name": "general_information"
  }
]