Research
Three questions organize my methodological work: which covariates an observational study should adjust for, how honestly the resulting uncertainty can be reported once the answer is contested or data-driven, and what measurement itself does to the data we analyse. The substantive programmes that follow use Japan's long-running panel surveys, administrative data, and newly digitized historical sources to study care, gender, place, policy, and the long shadow of past institutions over the life course.
Methodology
Causally disciplined multiverse analysis
Multiverse and specification-curve analyses estimate an effect across every plausible specification and report the distribution. But the specification space pools models that are causally incompatible: some control sets identify the declared effect, others condition on mediators or colliders and identify nothing of the kind. I develop a framework that encodes rival causal accounts as a small set of candidate directed acyclic graphs, restricts the multiverse to the specifications each graph licenses, and decomposes total multiverse variance into a within-graph component (sampling and specification noise) and a between-graph component (structural uncertainty). Reanalyses of canonical multiverses — the female-hurricanes controversy, the LaLonde job-training sample, and the union wage premium — show that pooled robustness metrics can manufacture both robustness and fragility. An open-source R package, dagmv, implements the framework; a panel-data extension is in progress.
Inference after data-driven covariate selection
The outcome-adaptive lasso is the most popular data-driven confounder selector in biostatistics, yet inference after it has remained naive. I study when uniformly valid confidence intervals are possible after selection rules that exclude instruments, propose a double-selection estimator with cross-fitted doubly robust estimation that attains validity and efficiency under an interpretable separation condition, and characterize the honest alternatives when that condition fails.
Estimand-specific covariate adjustment
Graphical theory tells us which adjustment set is most efficient for the average treatment effect. For the estimands applied work actually reports — effects on the treated, overlap-weighted effects — no such theory existed. I derive the adjustment-set calculus for the treated-population estimand and show that, unlike the ATE, its optimal adjustment set is not determined by the causal graph alone.
Measurement in panel surveys
Panel surveys assume that being measured does not change what is measured. Six decades of research on panel conditioning show that the assumption can fail, but the literature communicates through significance verdicts rather than distributions, and none of it tells an analyst what to believe about a panel that cannot estimate its own conditioning. I treat the problem as one of identification. On the theory side I ask what the level of the conditioning path can mean in a panel with staggered entry, where tenure, period, and cohort are linearly dependent, and what refreshment samples buy: which designs — survival matching, a symmetric variant of it, and a correction that uses a cohort's own unconditioned first-wave answers — identify the effect among persistent respondents, under what assumptions about attrition, and how those assumptions can be checked against one another; an R package, panelcond, implements the designs. On the empirical side, a full-questionnaire diagnosis of the nineteen-wave Japanese Life Course Panel Surveys asks whether repeated interviewing changes how people answer or what they answer, and how much of the published detection rate depends on survivor designs; a two-wave web experiment with a fresh-sample arm separates the mechanisms; and a meta-analysis of six decades of estimates converts the accumulated evidence into transportable bias priors for panels that have no refreshment sample of their own. Related work asks what nineteen annual measurements of a subjective-status item can separate — reliability from true-score stability — validates retrospective life histories against a prospective panel benchmark, bridges the class-identification item Japanese surveys have used since 1955 to the ten-rung ladder used internationally, and examines what the spatial unit at which “neighbourhood” is measured does to neighbourhood-effects estimates.
Substantive programmes
Care work, long-term care, and the life course
My longest-running substantive interest is the care economy in an ageing society: wages and turnover of care workers, the labour-market consequences of informal caregiving, and, in a current JSPS-funded project, households that combine childcare with elder care (“dual caregiving”). Much of this work draws on the Japanese Life Course Panel Surveys and on government statistics — the same panels whose measurement properties the project above examines.
Gender inequality in Japan: careers, authority, and pay disclosure
A second programme asks why Japan's gender gaps in pay, careers, and workplace authority have proved so immobile, and assembles evidence at three levels. At the level of individual lives, panel estimates link nineteen waves of the Japanese Life Course Panel Surveys to municipal childcare supply to ask how large mothers' long-run earnings penalty is, whether it varies with the childcare environment, and how it is graded by gender-role attitudes and the household division of labour. At the level of careers, retrospective work histories spanning five decennial waves of the Social Stratification and Social Mobility surveys ask how women's employment trajectories and their access to workplace authority have changed across sixty birth cohorts, and whether educational expansion and legal change moved either. At the level of firms, a panel of some 21,000 employers assembled from Japan's pay-gap disclosure platform and its web archive asks what a transparency mandate without sanctions produces. A methodological companion audits the retrospective histories on which the cohort work rests.
Stratification, life courses, and marriage markets
Japan's two national data systems — the decennial Social Stratification and Social Mobility surveys and the Japanese Life Course Panel Surveys, now nineteen consecutive annual waves — support descriptions of stratification at a scale previously out of reach. I use them to ask how employment careers and family formation are jointly produced: trajectory typologies built from nineteen-wave employment sequences, a marriage-cohort design spliced across five survey waves to locate which educational boundaries in the marriage market have softened, and an analysis of whether the shift from arranged introductions to matching apps changed who marries whom.
Residential immobility, place, and the life course
A third programme asks what staying does — to whom, and at whose expense — in a society where most adults never move. Nineteen annual waves of postal-code residential histories from the Japanese Life Course Panel Surveys are linked to small-area census deprivation indices, official land prices, and statutory flood and landslide maps, so that whole residential biographies, rather than isolated moves, become the unit of analysis. The programme asks what anchors people to a place — assets, age, family obligation, or the quality of the place itself — who leaves deprived or hazardous neighbourhoods and who cannot, where the returns to moving are largest and whether they accrue to the people able to move, and how the geographic anchoring of successor children changes when a parent dies. Papers from the programme are in preparation.
Natural experiments in Japanese social policy
Japanese social policy changes in ways that are unusually legible: a fundraising rule takes effect on a single date, a category of long-term-care bed is legislated out of existence, an admission threshold moves, a public hospital closes and its successor opens three kilometres away. Paired with administrative data covering every municipality, these discontinuities support designs whose identifying variation can be stated in a sentence. The organizing commitment of the programme is that inference discipline belongs to the design rather than to a robustness appendix: each study reports, next to its estimate, the effect size the design could actually have detected. One strand asks what regulation does to a competitive field, using Japan's hometown-tax programme — which moves more than a trillion yen a year between municipalities competing through reciprocal gifts — to study how municipalities responded to a ceiling on solicitation costs and how near-total convergence of practice arises in an organizational field. A second strand concerns health, care, and mortality in an ageing society: where people die when long-term-care beds are abolished, and what two decades of public-hospital consolidations and closures did to resident mortality once population ageing is accounted for. One methodological problem recurs across these designs and has become a paper of its own. Several of them compare a handful of treated prefectures or municipalities against many untreated ones, a situation in which cluster-robust standard errors over-reject badly and the standard remedy fails silently in the opposite direction; I calibrate what the available procedures do to size and power at the cluster counts applied work actually encounters, and use randomization inference throughout. Working papers and pre-analysis plans are posted as they are registered.
Long-run persistence and the construction of historical data
Some of what social policy can achieve today was settled long before the policy was written, and testing that claim usually requires sources that do not yet exist as data. Both halves of that sentence are part of the work. Imperial Ordinance No. 48 of 1887 barred Japanese prefectures from financing medical schools out of local taxes and eliminated most of them; a handful survived through idiosyncratic budgetary and administrative channels unrelated to their medical or fiscal strength. I ask whether that accident still shapes the geography of physician supply, how any such effect unfolded across a century of equalization policy and the 2004 liberalization of residency matching, and which mechanism — local training pipelines, amenities, or income — carries it. Establishing this meant building the sources: a prefectural physician census transcribed from the Imperial Statistical Yearbook of 1885, and six editions of the Nihon Isekiroku, a commercial medical directory digitized through the National Diet Library's OCR infrastructure, which yield 229,181 structured entries linked by license number into a panel of 47,084 physicians between 1925 and 1942. The register panel is being prepared for release as a documented dataset, with its parsing pipeline and an accuracy audit. Work now under way extends the panel further back, to the enrollment rosters of late-Edo private academies, and asks whether exposure to schools of Dutch learning left a trace in the medical geography of the century that followed.
Funding
- 2026–29 JSPS KAKENHI Grant-in-Aid for Scientific Research (C) 26K05332 (PI): dual caregiving across childcare and elder care.
- 2022–25 JSPS KAKENHI Early-Career Scientists 22K13525 (PI): diagnosis, impact, and correction of measurement error in social surveys.
- 2019–21 JSPS KAKENHI Early-Career Scientists 19K13907 (PI): care suppliers in a super-ageing society.
- 2017–18 JSPS KAKENHI Research Activity Start-up 17H06829 (PI): cooperative behaviour in medical institutions.
- 2018–23 Co-investigator, KAKENHI (A) 18H03630 (PI: M. Nakabayashi) and Challenging Research 18K18594 (PI: S. Fujihara).