Jörn Boehnke*

University of California, Davis

This version: September 2026

jb@ucdavis.edu

Abstract

I am an applied microeconomist who studies firms and markets and develops computational methods for extracting economic information from text, transactions, and high-frequency records. My current research focuses on language models, both as tools for economic measurement and as objects of study. I develop text-classification workflows that balance accuracy, computational cost, and human review, and methods that train small language models on their own elicited reasoning. I apply these tools to questions about pricing, labor demand, consumer search, and platform design. Across my work, I combine economic models, careful data construction, and machine learning to measure behavior that conventional datasets do not directly record and to understand when new computational methods materially improve economic analysis.

JEL codes: D83, L86, M31, C45, C55.

Keywords: online markets; machine learning; consumer search; large language models; price dispersion; unstructured data.

1Research

Bodoh-Creed, Aaron, Jörn Boehnke, and Brent Hickman. Forthcoming. “Market Frictions Versus Seller Signaling: Explaining Deviations from the Law of One Price.” Quantitative Marketing and Economics.

Abstract: Market Frictions Versus Seller Signaling: Explaining Deviations from the Law of One Price

Significant price dispersion for homogeneous goods in e-commerce is a well-documented puzzle; we find robust empirical evidence that the puzzle is roughly half as large as previously thought. We use an approach motivated by economic theory, employing a rich new dataset encompassing a wealth of unstructured data from 14 product categories, and flexible machine-learning methods. We quantify sellers’ marketing strategies through customization of the content, layout, and appearance of their product-listing web pages. Price variation systematically associated with these observable listing characteristics captures perceived listing heterogeneity rather than pricing noise generated by market frictions, allowing us to bound the role of market frictions in observed price dispersion. We find that seller-customized listing characteristics account for roughly half of observed cross-sectional price dispersion. These characteristics also predict buyer responses, with more information-rich listings associated with higher sale probabilities and shorter listing durations. Thus, a substantial portion of apparent deviations from the Law of One Price reflects economically meaningful differences in how sellers present otherwise physically identical products.

Boehnke, Jörn, Jacob Brophy, and Ashwin A Nair. 2026. “Self-Distilling Mathematical Reasoning in Small Language Models.” Revise & resubmit, Transactions on Machine Learning Research.

Abstract: Self-Distilling Mathematical Reasoning in Small Language Models

Small base language models (0.5B–3B parameters) often fail to produce structured, scoreable output in zero-shot mathematical reasoning, leaving reward- and preference-based optimization with little signal to learn from. Yet the same models reach 80–90% scoreability under few-shot prompting, so the reasoning capability is present but not expressed by the zero-shot policy. We propose Advantage-Weighted Direct Preference Optimization (AWDPO), a self-distillation method that trains models to reproduce their own few-shot reasoning behavior without exemplars at inference time. AWDPO weights preference updates by the reward gap between an elicited few-shot trace and the model’s own zero-shot attempt, and stabilizes training with a dynamically scaled likelihood anchor. On GSM8K with Qwen-2.5 base models, AWDPO recovers over 90% of the accuracy of fully supervised chain-of-thought fine-tuning using four chain-of-thought exemplars rather than 7,473 annotated traces, and the resulting models generalize zero-shot to SVAMP, ASDiv, and MATH-500 without additional training. Methods that do not train on elicited traces fare far worse, with DPO and GRPO reaching 0% at 0.5B and answer-only supervised fine-tuning never exceeding 14%. We then examine the source of these gains by comparing AWDPO against single-round self-distillation on few-shot generated traces from each base model, with parameterization held fixed. The simpler objective comes within 2.2 points at 0.5B and 1.5B in-domain under the matched-LoRA comparison and transfers better out-of-domain at those scales. What closes the cold-start gap is the elicitation, not the optimization objective applied on top of it.

Boehnke, Jörn, Pantelis Loupos, and Ying Gu. 2024. “Social drug dealing: how peer-to-peer fintech platforms have transformed illicit drug markets.” Annals of Operations Research, 335(2), 645-663.

Abstract: Social drug dealing: how peer-to-peer fintech platforms have transformed illicit drug markets

Digital platforms have revolutionized the way illegal drug trafficking is taking place. Modern drug dealers use social network platforms, such as Instagram and TikTok, as direct-to-consumer marketing tools. But apart from the marketing side, drug dealers also use fintech payment apps to engage in financial transactions with their clients. In this work, we leverage a large dataset from Venmo to investigate the digital money trail of drug dealers and the social networks they create. Using text and social network analytics, we identify two types of illicit users: mixed-activity participants and heavy drug traffickers and build a random forest classifier that accurately predicts both types of illicit nodes. We then investigate the social network structure of drug dealers on Venmo and find that heavy drug traffickers share similar network characteristics with previous literature findings on drug trafficking networks. However, mixed-activity participants exhibit different patterns of network structure characteristics, including a higher clustering coefficient, suggesting that they may be accessing multiple networks and bridging those networks through their illicit activities. Our findings highlight the importance of distinguishing between these two types of illicit users and provide law enforcement agencies with valuable insights that can aid in combating illegal drug transactions in digital payment apps.

Aravindakshan, Ashwin, Jörn Boehnke, Ehsan Gholami, and Ashutosh Nayak. 2022. “The impact of mask-wearing in mitigating the spread of COVID-19 during the early phases of the pandemic.” PLOS Global Public Health, 2(9), e0000954.

Abstract: The impact of mask-wearing in mitigating the spread of COVID-19 during the early phases of the pandemic

Masks have been widely recommended as a precaution against COVID-19 transmission. Several studies have shown the efficacy of masks at reducing droplet dispersion in lab settings. However, during the early phases of the pandemic, the usage of masks varied widely across countries. Using individual response data from the Imperial College London -- YouGov personal measures survey, this study investigates the effect of mask use within a country on the spread of COVID-19. The survey shows that mask-wearing exhibits substantial variations across countries and over time during the pandemic’s early phase. We use a reduced form econometric model to relate population-wide variation in mask-wearing to the growth rate of confirmed COVID-19 cases. The results indicate that mask-wearing plays an important role in mitigating the spread of COVID-19. Widespread mask-wearing within a country associates with an expected 7% (95% CI: 3.94%-9.99%) decline in the growth rate of daily active cases of COVID-19 in the country. This daily decline equates to an expected 88.5% drop in daily active cases over a 30-day period when compared to zero percent mask-wearing, all else held equal. The decline in daily growth rate due to the combined effect of mask-wearing, reduced outdoor mobility, and non-pharmaceutical interventions averages 28.1% (95% CI: 24.2%-32%).

Boehnke, Jörn, and Victor Gay. 2022. “The Missing Men: World War I and Female Labor Force Participation.” Journal of Human Resources, 57(4), 1209-1241.

Abstract: The Missing Men: World War I and Female Labor Force Participation

Using spatial variation in World War I military fatalities in France, we show that the scarcity of men due to the war generated an upward shift in female labor force participation that persisted throughout the interwar period. Available data suggest that increased female labor supply accounts for this result. In particular, deteriorated marriage market conditions for single women and negative income shocks to war widows induced many of these women to enter the labor force after the war. In contrast, demand factors such as substitution toward female labor to compensate for the scarcity of male labor were of second-order importance.

Bracci, Alberto, Jörn Boehnke, Abeer ElBahrawy, Nicola Perra, Alexander Teytelboym, and Andrea Baronchelli. 2022. “Macroscopic properties of buyer-seller networks in online marketplaces.” PNAS Nexus, 1(4), pgac201.

Abstract: Macroscopic properties of buyer-seller networks in online marketplaces

Online marketplaces are the main engines of legal and illegal e-commerce, yet their empirical properties are poorly understood due to the absence of large-scale data. We analyze two comprehensive datasets containing 245M transactions (16B USD) that took place on online marketplaces between 2010 and 2021, covering 28 dark web marketplaces, i.e., unregulated markets whose main currency is Bitcoin, and 144 product markets of one popular regulated e-commerce platform. We show that transactions in online marketplaces exhibit strikingly similar patterns despite significant differences in language, lifetimes, products, regulation, and technology. Specifically, we find remarkable regularities in the distributions of transaction amounts, number of transactions, inter-event times and time between first and last transactions. We show that buyer behavior is affected by the memory of past interactions and use this insight to propose a model of network formation reproducing our main empirical observations. Our findings have implications for understanding market power on online marketplaces as well as inter-marketplace competition, and provide empirical foundation for theoretical economic models of online marketplaces.

Bodoh-Creed, Aaron L., Jörn Boehnke, and Brent R. Hickman. 2021. “How Efficient are Decentralized Auction Platforms?The Review of Economic Studies, 88(1), 91-125.

Abstract: How Efficient are Decentralized Auction Platforms?

We model a decentralized, dynamic auction market platform in which a continuum of buyers and sellers participate in simultaneous, single-unit auctions each period. Our model accounts for the endogenous entry of agents and the impact of intertemporal optimization on bids. We estimate the structural primitives of our model using Kindle sales on eBay. We find that just over one third of Kindle auctions on eBay result in an inefficient allocation with deadweight loss amounting to 14% of total possible market surplus. We also find that partial centralization -- for example, running half as many 2-unit, uniform-price auctions each day -- would eliminate a large fraction of the inefficiency, but yield lower seller revenues. Our results also highlight the importance of understanding platform composition effects -- selection of agents into the market -- in assessing the implications of market redesign. We also prove that the equilibrium of our model with a continuum of buyers and sellers is an approximate equilibrium of the analogous model with a finite number of agents.

Aravindakshan, Ashwin, Jörn Boehnke, Ehsan Gholami, and Ashutosh Nayak. 2020. “Preparing for a future COVID-19 wave: insights and limitations from a data-driven evaluation of non-pharmaceutical interventions in Germany.” Scientific Reports, 10, 20084.

Abstract: Preparing for a future COVID-19 wave: insights and limitations from a data-driven evaluation of non-pharmaceutical interventions in Germany

To contain the COVID-19 pandemic, governments introduced strict Non-Pharmaceutical Interventions (NPI) that restricted movement, public gatherings, national and international travel, and shut down large parts of the economy. Yet, the impact of the enforcement and subsequent loosening of these policies on the spread of COVID-19 is not well understood. Accordingly, we measure the impact of NPIs on mitigating disease spread by exploiting the spatio-temporal variations in policy measures across the 16 states of Germany. While this quasi-experiment does not allow for causal identification, each policy’s effect on reducing disease spread provides meaningful insights. We adapt the Susceptible-Exposed-Infected-Recovered (SEIR) model for disease propagation to include data on daily confirmed cases, interstate movement, and social distancing. By combining the model with measures of policy contributions on mobility reduction, we forecast scenarios for relaxing various types of NPIs. Our model finds that in Germany policies that mandated contact restrictions (e.g., movement in public space limited to two persons or people co-living), closure of educational institutions (e.g., schools), and retail outlet closures are associated with the sharpest drops in movement within and across states. Contact restrictions appear to be most effective at lowering COVID-19 cases, while border closures appear to have only minimal effects at mitigating the spread of the disease, even though cross-border travel might have played a role in seeding the disease in the population. We believe that a deeper understanding of the policy effects on mitigating the spread of COVID-19 allows a more accurate forecast of disease spread when NPIs are partially loosened and gives policymakers better data for making informed decisions.

Bhargava, Hemant K., Olivier Rubel, Elizabeth J. Altman, Ramnik Arora, Jörn Boehnke, Kaitlin Daniels, Timothy Derdenger, Bryan Kirschner, Darin LaFramboise, Pantelis Loupos, Geoffrey Parker, and Adithya Pattabhiramaiah. 2020. “Platform data strategy.” Marketing Letters, 31(4), 323-334.

Abstract: Platform data strategy

Platforms create value by enabling interactions between consumers and external producers through infrastructures and rules. We define platform data strategy to encompass all data-related rules undertaken by platforms to foster competitive advantage over the long-term. Platform firms face growing pressure to increase accountability for how they use data; yet, an explicit treatment of platforms’ data strategies and a systematic discussion of forces influencing such data-related choices is absent in the academic literature. We articulate how a platform’s data strategy varies based on platform type and business circumstances. Given the interdependencies within a platform’s ecosystem, its data strategy must balance incentives of all stakeholders. Besides discussing these topics, the paper identifies promising research opportunities in platform data strategy to better inform future academic research, strategic decision-making, and regulatory analysis.

Boehnke, Jörn, and Pietro Bonaldi. 2019. “Synthetic Regression Discontinuity: Estimating Treatment Effects using Machine Learning.” Peer-reviewed manuscript, NeurIPS 2019 workshop “Do the right thing”: machine learning and causal inference for improved decision making.

Kho, Abel N., John P. Cashy, Kathryn L. Jackson, Adam R. Pah, Satyender Goel, Jörn Boehnke, John Eric Humphries, Scott Duke Kominers, Bala N. Hota, Shannon A. Sims, Bradley A. Malin, Dustin D. French, Theresa L. Walunas, David O. Meltzer, Erin O. Kaleba, Roderick C. Jones, and William L. Galanter. 2015. “Design and implementation of a privacy preserving electronic health record linkage tool in Chicago.” Journal of the American Medical Informatics Association, 22(5), 1072-1080.

Abstract: Design and implementation of a privacy preserving electronic health record linkage tool in Chicago

Objective. To design and implement a tool that creates a secure, privacy preserving linkage of electronic health record (EHR) data across multiple sites in a large metropolitan area in the United States (Chicago, IL), for use in clinical research.

Methods. The authors developed and distributed a software application that performs standardized data cleaning, preprocessing, and hashing of patient identifiers to remove all protected health information. The application creates seeded hash code combinations of patient identifiers using a Health Insurance Portability and Accountability Act compliant SHA-512 algorithm that minimizes re-identification risk. The authors subsequently linked individual records using a central honest broker with an algorithm that assigns weights to hash combinations in order to generate high specificity matches.

Results. The software application successfully linked and de-duplicated 7 million records across 6 institutions, resulting in a cohort of 5 million unique records. Using a manually reconciled set of 11 292 patients as a gold standard, the software achieved a sensitivity of 96% and a specificity of 100%, with a majority of the missed matches accounted for by patients with both a missing social security number and last name change. Using 3 disease examples, it is demonstrated that the software can reduce duplication of patient records across sites by as much as 28%.

Conclusions. Software that standardizes the assignment of a unique seeded hash identifier merged through an agreed upon third-party honest broker can enable large-scale secure linkage of EHR data for epidemiologic and public health research. The software algorithm can improve future epidemiologic research by providing more comprehensive data given that patients may make use of multiple healthcare systems.

2Teaching

My teaching at UC Davis spans the Master of Science in Business Analytics (MSBA), MBA, and Online MBA programs.

Professor of the Year, MSBA Program, 2021, 2023, and 2026.

Machine Learning & Artificial Intelligence
MSBA, BAX-452. This course spans regression and regularization, tree-based methods, neural networks, language models, reinforcement learning, model selection, and evaluation.
Data Design & Representation
MSBA, BAX-422. This course develops methods for extracting, representing, and structuring unstructured data for analysis.
AI and Business Innovation: A Journey from Linear Regression to LLMs
MBA, MGV-490A. This course connects regression and causal reasoning to predictive modeling, neural networks, and large language models.
Agentic AI: From Models to Systems
MBA, MGV-490B. This course builds agentic systems that coordinate models, process news, use external tools, and execute decisions through MCP.
Data Wrangling
MBA, MGT/B-435. This course turns heterogeneous web sources into reproducible, analysis-ready datasets through automated acquisition and processing.
Business of the Future
Wine Executive Program, University of California, Davis. This course examines how AI and data shape customer strategy, profitability, and managerial decision-making.
Economic Research Experience for Undergraduates
BA, Becker Friedman Institute, University of Chicago. This course introduces empirical economic research through programming, data construction, measurement, and analysis.
New Tools for Acquisition / Analysis of Internet Data
PhD, Harvard University and NBER. This course develops methods for turning internet information into usable data for empirical economics.

3Curriculum Vitae

I received my Ph.D. in Economics from the University of Chicago and hold degrees in mathematics and physics from Leipzig University. From 2015 to 2024, I held research appointments at Harvard University’s Center of Mathematical Sciences and Applications, first as a Postdoctoral Research Fellow and later as an Associate Research Fellow. Further details are in my curriculum vitae.

AAppendix: Google Scholar

Bibliographic records and citations to my work are available on Google Scholar.

Contents