US MEPS Survey — Survey Overview 美国 MEPS 调查 — 概览
HC-251 · 2023Before exploring 1,366 variables, a few minutes on where this data comes from and how it's collected — especially the Panel / Round design, since notation like "R3/1" or "REFPRS31" shows up throughout the explorer and doesn't mean anything without this context.
在探索这 1,366 个变量之前,先花几分钟了解一下这份数据的来源和采集方式——尤其是 Panel(批次)/ Round(轮次)的设计,因为像 "R3/1" 或 "REFPRS31" 这样的记号会贯穿整个探索工具, 不了解背景的话根本看不懂。
The Medical Expenditure Panel Survey (MEPS) is run by the Agency for Healthcare Research and Quality (AHRQ) and has collected data every year since 1996. It produces the U.S. government's core estimates of health care use, spending, sources of payment, and insurance coverage for the civilian noninstitutionalized population.
MEPS has two linked parts. The Household Component (HC) interviews a nationally representative sample of households directly — this is the source for everything in this explorer. The Medical Provider Component (MPC) separately contacts the pharmacies, doctors, and hospitals that HC respondents named, and verifies or fills in payment and drug details the household couldn't accurately report themselves — it isn't designed to produce its own national estimates, just to improve the HC data.
Sampled by household, recorded by person. The household is only the entry point: MEPS
samples dwelling units to reach people, but the single household respondent then answers
questions about every household member individually — this person's age, this person's insurance,
this person's doctor visits. Since the facts collected are inherently person-specific, AHRQ builds the
Consolidated PUF with one row per person, not one per household — a 4-person household contributes 4
person-records, all from the same interview. Household/family ID variables (DUID,
FAMID23, RUSIZE23, etc.) are still on the file, so you can roll person-records
back up to the family level if you need to.
Source: AHRQ, 2023 MEPS HC-251 Full-Year Consolidated File documentation, §B 1.0 ("All data for a sampled household are reported by a single household respondent... Data can be analyzed at the person, the family, or the event level," p. B-1). h251doc.pdf →
This explorer covers the 2023 Full-Year Consolidated File (HC-251) — one person per record, 18,640 inscope respondents, 1,374 variables spanning demographics, income, health status, employment, insurance, and health care utilization/expenditure.
医疗支出面板调查(Medical Expenditure Panel Survey,MEPS)由美国医疗保健研究与质量局 (AHRQ)负责运营,自1996年起每年持续采集数据。它是美国政府关于医疗使用、医疗支出、支付来源,以及非机构化 平民人口医疗保险覆盖情况的核心官方估计数据来源。
MEPS 由两个相互关联的部分组成。家庭部分(Household Component,HC)直接访谈一个具有全国 代表性的家庭样本——本探索工具里的所有数据都来自这里。医疗服务提供方部分(Medical Provider Component,MPC)则单独联系 HC 受访者提到的药房、医生和医院,核实或补充家庭自己说不清楚的支付和 用药细节——它本身并不用来产出独立的全国估计,只是用来改善 HC 数据的质量。
按家庭抽样,按人记录。家庭只是"入口":MEPS 抽样的单位是住宅(dwelling unit),
目的是借此触达具体的人——但那位家庭受访者接下来要逐个回答家里每一个成员的具体情况:这个人的
年龄、这个人的保险、这个人的就医记录。因为采集到的事实天然就是"某个具体人"的事实,所以 AHRQ 最终建的
合并文件是一人一条记录,不是一户一条记录——一户4口之家会在文件里产生4条人物记录,全部来自同一次访谈。
文件里仍然保留了家庭/户的标识变量(DUID、FAMID23、RUSIZE23 等),
所以有需要的话依然可以把人物记录重新汇总回家庭层面。
资料来源:AHRQ《2023年 MEPS HC-251 年度合并文件》官方文档 §B 1.0 ("All data for a sampled household are reported by a single household respondent... Data can be analyzed at the person, the family, or the event level.",第 B-1 页)。 h251doc.pdf →
本探索工具覆盖的是2023年度完整合并文件(Full-Year Consolidated File,HC-251)——每条记录 对应一个人,共 18,640 名符合调查范围的受访者,1,374 个变量,涵盖人口特征、 收入、健康状况、就业、保险以及医疗使用/支出等方面。
Each year AHRQ recruits a new Panel — a fresh cohort of households, drawn as a subsample of households that answered the prior year's National Health Interview Survey (NHIS). For the two panels behind this file, Panel 27 sampled 9,700 households (from the 2021 NHIS) and Panel 28 sampled 9,800 households (from the 2022 NHIS). Every household in a panel is then interviewed 5 times ("Rounds") about 5 months apart, covering roughly 2.5 calendar years total, so each panel's data collection straddles two calendar years.
Because a new panel starts every year while the previous one is still mid-way through its 5 rounds, two panels are always running at once, offset by about a year. Any single calendar year's data — like 2023 — is stitched together from the second half of one panel and the first half of the next. §04 below walks through exactly how those ~19,500 sampled households turn into this file's 18,463 final person-records.
Source: h251doc.pdf §3.1.1, "Panel 27/28 Household Sample Size."
AHRQ 每年都会招募一个新的 Panel(批次)——从上一年参加过全国健康访谈调查(NHIS)的家庭中 抽取子样本组成。本文件背后的两个批次里,Panel 27 从2021年 NHIS 中抽了 9,700 户,Panel 28 从2022年 NHIS 中抽了 9,800 户。批次内的每户家庭之后会被访谈5次(称为 "Round/轮次"),每次间隔约5个月,整个过程跨度约2.5个日历年,所以每个批次的数据采集都会横跨两个 日历年。
由于新批次每年都会启动,而上一个批次的5轮访谈还没走完,所以任何时候都同时有两个批次在进行, 彼此错开大约一年。任何单一日历年的数据——比如2023年——都是由一个批次的后半段和下一个批次的 前半段拼接而成的。这约19,500户抽样家庭具体是怎么变成本文件最终18,463条人物记录的,见下面第04节。
资料来源:h251doc.pdf §3.1.1,"Panel 27/28 Household Sample Size"。
That's the whole trick behind the "R3/1", "R4/2", "R5/3" notation and variable suffixes like
REFPRS31: the first digit is Panel 27's round number, the second is Panel 28's round number
for that same time window — Round 3 of Panel 27 and Round 1 of Panel 28 happened concurrently, so
they're referred to together as "R3/1" (variable suffix 31). Likewise R4/2 → 42,
R5/3 → 53. A variable ending in plain 23 (no round pairing) means "status as of
December 31, 2023," combining whichever round's data applies for that person.
Source: AHRQ, 2023 MEPS HC-251 Full-Year Consolidated File documentation, §B/C 1.0–2.4 and the "MEPS Panel Design: Data Reference Periods" chart (p. C-2). h251doc.pdf →
"R3/1"、"R4/2"、"R5/3" 这种记号,以及像 REFPRS31 这样的变量后缀,秘密就在这里:第一位数字是
Panel 27 的轮次号,第二位是同一时间窗口里 Panel 28 的轮次号——Panel 27 的第3轮和 Panel 28 的第1轮
是同时进行的,所以合称 "R3/1"(变量后缀 31)。同理 R4/2 → 42,R5/3 →
53。如果变量结尾是单纯的 23(没有轮次配对),代表"截至2023年12月31日的状态",
具体取哪一轮的数据取决于这个人当时所处的轮次。
资料来源:AHRQ《2023年 MEPS HC-251 年度合并文件》官方文档 §B/C 1.0–2.4,及 "MEPS Panel Design: Data Reference Periods" 图表(第 C-2 页)。 h251doc.pdf →
Each household has one respondent — usually the person most knowledgeable about the family's health and health care — who answers for everyone in the household across all 5 rounds, via computer-assisted personal interviewing (CAPI). AHRQ distinguishes a few respondent roles that show up repeatedly in the "Survey Administration" variables:
每户家庭都有一位受访者(respondent)——通常是家里对健康和医疗情况最了解的人——在全部5轮 访谈中代表全家作答,访谈方式是计算机辅助面对面访谈(CAPI)。AHRQ 在 "Survey Administration" 这组变量里 反复区分了几种受访者角色:
Combining two panels means combining two separate response funnels — households sampled, then whittled down round by round as some don't respond. AHRQ's Table 19 gives the exact 2023 numbers:
合并两个批次,意味着要合并两条各自独立的应答漏斗——先抽样出户数,再逐轮因为无应答而递减。AHRQ 官方 Table 19 给出了2023年的精确数字:
4,262 + 5,143 = 9,405 households (technically "Reporting Units," which can be more
granular than a physical household) completed every round required of them — an overall unweighted
response rate of 26.1% (Panel 27: 24.1%, Panel 28: 27.4%, combined by weighting them 0.40
/ 0.60). From RU to person: those 9,405 completing households contained 18,463 people who
ended up with a positive final weight (PERWT23F) — averaging just under 2 people per
household, which is why the final person count lands close to the household counts above rather than far
above them.
Source: h251doc.pdf §3.2, Table 19.
4,262 + 5,143 = 9,405 户(严格来说是"申报单位 RU",可能比物理意义上的一户更细)完成了各自
所需的全部轮次——整体未加权应答率为26.1%(Panel 27为24.1%,Panel 28为27.4%,按0.40/0.60
加权合并)。从 RU 到人:这9,405户完成访谈的家庭里,共有 18,463 人最终拥有正的权重
(PERWT23F)——平均每户略少于2人,这就是为什么最终的人数会落在和户数相近的量级,而不是远
高于户数。
资料来源:h251doc.pdf §3.2,Table 19。
The year-over-year decline below isn't just "a different panel happened to be smaller" — AHRQ documents a specific cause. In its own words: "MEPS has been substantially affected by the pandemic... One effect of the pandemic is the significantly lower response rates... analysts... should continue to exercise caution when interpreting estimates and assessing analyses, especially for data collected from 2020 through 2022. This includes the comparison of such estimates to those of other years and corresponding trend analyses." AHRQ's sample design modifications made to compensate ended in 2022, but response rates themselves haven't recovered to pre-pandemic levels — this is a broader trend across nearly all U.S. government surveys in this period, not unique to MEPS.
Source: h251doc.pdf §3.1.2, "Discussion of Pandemic Effects on Quality of MEPS Data."
下面这个逐年下降的趋势,不只是"刚好换了个更小的批次"这么简单——AHRQ 自己记录了具体原因。原文是: "MEPS has been substantially affected by the pandemic... One effect of the pandemic is the significantly lower response rates... analysts... should continue to exercise caution when interpreting estimates and assessing analyses, especially for data collected from 2020 through 2022. This includes the comparison of such estimates to those of other years and corresponding trend analyses." ("MEPS 受到疫情的显著影响……疫情带来的影响之一就是应答率明显走低……分析人员在解读估计值时应格外谨慎, 尤其是2020至2022年采集的数据……这也包括与其他年份做比较、做趋势分析的情况。")AHRQ 为应对疫情做的抽样 设计调整在2022年就结束了,但应答率本身并没有恢复到疫情前的水平——这是同期几乎所有美国政府调查都遇到的 更普遍的趋势,并非 MEPS 独有。
资料来源:h251doc.pdf §3.1.2,"Discussion of Pandemic Effects on Quality of MEPS Data"。
This explorer's 18,640 figure is the inscope-respondent count before the final
person-weight cut; 18,463 of those have a positive final weight
(PERWT23F) and are the set AHRQ says to use for U.S.-population-level estimates. All
statistics in this explorer are unweighted — raw respondent counts, not projected
national totals. Weighting exists specifically to correct for the panel design's uneven sampling and
non-response, so treat any percentage here as "share of respondents," not "share of the U.S. population."
Source: h251doc.pdf §C 2.0.
本探索工具里的 18,640 这个数字,是在做最终人物权重筛选之前的"调查范围内受访者"计数;其中
18,463 人拥有正的最终权重(PERWT23F),这是 AHRQ 官方建议用来做全美人口层面
估计的样本集合。本探索工具里的所有统计都是未加权的——是原始受访者计数,不是推算出的全国
总量。加权的存在正是为了修正批次设计里不均匀的抽样和无应答问题,所以这里出现的任何百分比,都应该理解为
"受访者中的占比",而不是"美国全体人口中的占比"。
资料来源:h251doc.pdf §C 2.0。