Digital platforms have revolutionized the way illegal drug trafficking is taking place. Modern drug dealers use social network platforms, such as Instagram and TikTok, as direct-to-consumer marketing tools. But apart from the marketing side, drug dealers also use fintech payment apps to engage in financial transactions with their clients. In this work, we leverage a large dataset from Venmo to investigate the digital money trail of drug dealers and the social networks they create. Using text and social network analytics, we identify two types of illicit users: mixed-activity participants and heavy drug traffickers and build a random forest classifier that accurately predicts both types of illicit nodes. We then investigate the social network structure of drug dealers on Venmo and find that heavy drug traffickers share similar network characteristics with previous literature findings on drug trafficking networks. However, mixed-activity participants exhibit different patterns of network structure characteristics, including a higher clustering coefficient, suggesting that they may be accessing multiple networks and bridging those networks through their illicit activities. Our findings highlight the importance of distinguishing between these two types of illicit users and provide law enforcement agencies with valuable insights that can aid in combating illegal drug transactions in digital payment apps.
1Research
1.1Publications
Masks have been widely recommended as a precaution against COVID-19 transmission. Several studies have shown the efficacy of masks at reducing droplet dispersion in lab settings. However, during the early phases of the pandemic, the usage of masks varied widely across countries. Using individual response data from the Imperial College London—YouGov personal measures survey, this study investigates the effect of mask use within a country on the spread of COVID-19. The survey shows that mask-wearing exhibits substantial variations across countries and over time during the pandemic’s early phase. We use a reduced form econometric model to relate population-wide variation in mask-wearing to the growth rate of confirmed COVID-19 cases. The results indicate that mask-wearing plays an important role in mitigating the spread of COVID-19. Widespread mask-wearing within a country associates with an expected 7% (95% CI: 3.94%–9.99%) decline in the growth rate of daily active cases of COVID-19 in the country. This daily decline equates to an expected 88.5% drop in daily active cases over a 30-day period when compared to zero percent mask-wearing, all else held equal. The decline in daily growth rate due to the combined effect of mask-wearing, reduced outdoor mobility, and non-pharmaceutical interventions averages 28.1% (95% CI: 24.2%–32%).
Using spatial variation in World War I military fatalities in France, we show that the scarcity of men due to the war generated an upward shift in female labor force participation that persisted throughout the interwar period. Available data suggest that increased female labor supply accounts for this result. In particular, deteriorated marriage market conditions for single women and negative income shocks to war widows induced many of these women to enter the labor force after the war. In contrast, demand factors such as substitution toward female labor to compensate for the scarcity of male labor were of second-order importance.
Online marketplaces are the main engines of legal and illegal e-commerce, yet their empirical properties are poorly understood due to the absence of large-scale data. We analyze two comprehensive datasets containing 245M transactions (16B USD) that took place on online marketplaces between 2010 and 2021, covering 28 dark web marketplaces, i.e., unregulated markets whose main currency is Bitcoin, and 144 product markets of one popular regulated e-commerce platform. We show that transactions in online marketplaces exhibit strikingly similar patterns despite significant differences in language, lifetimes, products, regulation, and technology. Specifically, we find remarkable regularities in the distributions of transaction amounts, number of transactions, inter-event times and time between first and last transactions. We show that buyer behavior is affected by the memory of past interactions and use this insight to propose a model of network formation reproducing our main empirical observations. Our findings have implications for understanding market power on online marketplaces as well as inter-marketplace competition, and provide empirical foundation for theoretical economic models of online marketplaces.
We model a decentralized, dynamic auction market platform in which a continuum of buyers and sellers participate in simultaneous, single-unit auctions each period. Our model accounts for the endogenous entry of agents and the impact of intertemporal optimization on bids. We estimate the structural primitives of our model using Kindle sales on eBay. We find that just over one third of Kindle auctions on eBay result in an inefficient allocation with deadweight loss amounting to 14% of total possible market surplus. We also find that partial centralization—for example, running half as many 2-unit, uniform-price auctions each day—would eliminate a large fraction of the inefficiency, but yield lower seller revenues. Our results also highlight the importance of understanding platform composition effects—selection of agents into the market—in assessing the implications of market redesign. We also prove that the equilibrium of our model with a continuum of buyers and sellers is an approximate equilibrium of the analogous model with a finite number of agents.
To contain the COVID-19 pandemic, governments introduced strict Non-Pharmaceutical Interventions (NPI) that restricted movement, public gatherings, national and international travel, and shut down large parts of the economy. Yet, the impact of the enforcement and subsequent loosening of these policies on the spread of COVID-19 is not well understood. Accordingly, we measure the impact of NPIs on mitigating disease spread by exploiting the spatio-temporal variations in policy measures across the 16 states of Germany. While this quasi-experiment does not allow for causal identification, each policy’s effect on reducing disease spread provides meaningful insights. We adapt the Susceptible-Exposed-Infected-Recovered (SEIR) model for disease propagation to include data on daily confirmed cases, interstate movement, and social distancing. By combining the model with measures of policy contributions on mobility reduction, we forecast scenarios for relaxing various types of NPIs. Our model finds that in Germany policies that mandated contact restrictions (e.g., movement in public space limited to two persons or people co-living), closure of educational institutions (e.g., schools), and retail outlet closures are associated with the sharpest drops in movement within and across states. Contact restrictions appear to be most effective at lowering COVID-19 cases, while border closures appear to have only minimal effects at mitigating the spread of the disease, even though cross-border travel might have played a role in seeding the disease in the population. We believe that a deeper understanding of the policy effects on mitigating the spread of COVID-19 allows a more accurate forecast of disease spread when NPIs are partially loosened and gives policymakers better data for making informed decisions.
Platforms create value by enabling interactions between consumers and external producers through infrastructures and rules. We define platform data strategy to encompass all data-related rules undertaken by platforms to foster competitive advantage over the long-term. Platform firms face growing pressure to increase accountability for how they use data; yet, an explicit treatment of platforms’ data strategies and a systematic discussion of forces influencing such data-related choices is absent in the academic literature. We articulate how a platform’s data strategy varies based on platform type and business circumstances. Given the interdependencies within a platform’s ecosystem, its data strategy must balance incentives of all stakeholders. Besides discussing these topics, the paper identifies promising research opportunities in platform data strategy to better inform future academic research, strategic decision-making, and regulatory analysis.
In the standard regression discontinuity setting, treatment assignment is based on whether a unit’s observable score (running variable) crosses a known threshold. We propose a two-stage method to estimate the treatment effect when the score is unobservable to the econometrician while the treatment status is known for all units. In the first stage, we use a statistical model to predict a unit’s treatment status based on a continuous synthetic score. In the second stage, we apply a regression discontinuity design using the predicted synthetic score as the running variable to estimate the treatment effect on an outcome of interest. We establish conditions under which the method identifies the local treatment effect for a unit at the threshold of the unobservable score, the same parameter that a standard regression discontinuity design with known score would identify. We also examine the properties of the estimator using simulations, and propose the use of machine learning algorithms to achieve high prediction accuracy. Finally, we apply the method to measure the effect of an investment grade rating on corporate bond prices by any of the three largest credit ratings agencies. We find an average 1% increase in the prices of corporate bonds that received an investment grade as opposed to a non-investment grade rating.
Objective. To design and implement a tool that creates a secure, privacy preserving linkage of electronic health record (EHR) data across multiple sites in a large metropolitan area in the United States (Chicago, IL), for use in clinical research.
Methods. The authors developed and distributed a software application that performs standardized data cleaning, preprocessing, and hashing of patient identifiers to remove all protected health information. The application creates seeded hash code combinations of patient identifiers using a Health Insurance Portability and Accountability Act compliant SHA-512 algorithm that minimizes re-identification risk. The authors subsequently linked individual records using a central honest broker with an algorithm that assigns weights to hash combinations in order to generate high specificity matches.
Results. The software application successfully linked and de-duplicated 7 million records across 6 institutions, resulting in a cohort of 5 million unique records. Using a manually reconciled set of 11 292 patients as a gold standard, the software achieved a sensitivity of 96% and a specificity of 100%, with a majority of the missed matches accounted for by patients with both a missing social security number and last name change. Using 3 disease examples, it is demonstrated that the software can reduce duplication of patient records across sites by as much as 28%.
Conclusions. Software that standardizes the assignment of a unique seeded hash identifier merged through an agreed upon third-party honest broker can enable large-scale secure linkage of EHR data for epidemiologic and public health research. The software algorithm can improve future epidemiologic research by providing more comprehensive data given that patients may make use of multiple healthcare systems.
1.2Working Papers
Significant price dispersion for homogeneous goods in e-commerce is a well-documented puzzle; we find robust empirical evidence that the puzzle is roughly half as large as previously thought. We use an approach motivated by economic theory, employing a rich new dataset encompassing a wealth of unstructured data from 14 product categories, and flexible machine-learning methods. We quantify sellers’ marketing strategies through customization of the content, layout, and appearance of their product-listing web pages. Price variation systematically associated with these observable listing characteristics captures perceived listing heterogeneity rather than pricing noise generated by market frictions, allowing us to bound the role of market frictions in observed price dispersion. We find that seller-customized listing characteristics account for roughly half of observed cross-sectional price dispersion. These characteristics also predict buyer responses, with more information-rich listings associated with higher sale probabilities and shorter listing durations. Thus, a substantial portion of apparent deviations from the Law of One Price reflects economically meaningful differences in how sellers present otherwise physically identical products.
Small base language models (0.5B–3B parameters) often fail to produce structured, scoreable output in zero-shot mathematical reasoning, leaving reward- and preference-based optimization with little signal to learn from. Yet the same models reach 80–90% scoreability under few-shot prompting, so the reasoning capability is present but not expressed by the zero-shot policy. We propose Advantage-Weighted Direct Preference Optimization (AWDPO), a self-distillation method that trains models to reproduce their own few-shot reasoning behavior without exemplars at inference time. AWDPO weights preference updates by the reward gap between an elicited few-shot trace and the model’s own zero-shot attempt, and stabilizes training with a dynamically scaled likelihood anchor. On GSM8K with Qwen-2.5 base models, AWDPO recovers over 90% of the accuracy of fully supervised chain-of-thought fine-tuning using four chain-of-thought exemplars rather than 7,473 annotated traces, and the resulting models generalize zero-shot to SVAMP, ASDiv, and MATH-500 without additional training. Methods that do not train on elicited traces fare far worse, with DPO and GRPO reaching 0% at 0.5B and answer-only supervised fine-tuning never exceeding 14%. We then examine the source of these gains by comparing AWDPO against single-round self-distillation on few-shot generated traces from each base model, with parameterization held fixed. The simpler objective comes within 2.2 points at 0.5B and 1.5B in-domain under the matched-LoRA comparison and transfers better out-of-domain at those scales. What closes the cold-start gap is the elicitation, not the optimization objective applied on top of it.
A retail chain must decide when to update prices and whether one rule can serve every location. We show how a shared rule can retain local timing without separate schedules. A fixed-time schedule permits increases at specified times. A gap trigger raises the price when the reference price exceeds the posted price by more than a threshold. Using 923 million timestamped fuel-price reports, we select all settings using 2024 data and test how closely the rules reproduce the 2025 diesel-price paths of 2,000 German stations, with a daily limit on increases and unrestricted cuts. Locally tuned schedules and triggers have nearly the same tracking error, but sharing their settings has very different effects. With one permitted increase per day, sharing adds 0.642 euro cents per liter to schedule error but only 0.097 to trigger error, a difference of 0.545. One shared trigger recovers 84 percent of the error reduction from replacing one shared schedule with 2,000 local schedules. Matching the shared trigger to within 0.10 cents per liter takes two schedule templates and station-specific assignments. The trigger’s sharing advantage persists within most brand networks. Two features explain the result: many stations can use the same threshold, and local price gaps determine when the trigger acts. With four permitted increases, a shared schedule that can wait for a positive gap outperforms 2,000 local schedules. Delayed information or limits that also count cuts can instead favor a fixed-time schedule over a trigger. A network can therefore share its update rule without imposing common update times.
Three months after Spotify let artists pitch songs directly to its playlist editors in 2018, more than 10,000 artists had gained their first editorial placement. But royalties follow recordings, not artists, and an artist who self-releases one song can release the next through a label. We ask how much of the music promoted for independent artists is independently released. We link 2016–2019 placements on Spotify’s editorial and major-label playlists to follower counts, 2026 catalog credits, and later releases. Spotify’s own playlists supply over nine-tenths of self-releasing artists’ potential playlist audience. Yet after pitching opened, the catalog classifies only about half of these artists’ editorial placements as self-released, and one fifth on the majors’ playlists. Among artists featured on both, the gap comes almost entirely from how often each is featured. The majors’ playlists also predict an artist’s next major label: of 70 self-releasing artists they featured who released with a major within a year, 54.3% released with the playlist’s owner, against 31.8% expected from the mix of firms, a pattern that predates direct pitching. Classifying placements by artist rather than recording thus overstates self-released placements about twofold on Spotify’s playlists and fivefold on the majors’.
Organizations increasingly delegate operational work to large language models, including the classification of unstructured text such as clinical notes, regulatory filings, and corporate communications. This paper develops empirically grounded guidelines for that practice and proposes a decision framework that treats classification-workflow design as an operational decision, addressing an open question in the use of large language models for classification: which cases should the model decide, which should be routed to human review, and at what cost? In the framework, a decision maker trades accuracy on resolved tasks against the cost of unresolved tasks and the computational cost of repeated model runs. We evaluate two primary configuration levers—self-reported confidence thresholds and redundancy (repeated runs of the same task)—in a controlled computational experiment spanning two applied classification tasks, corporate-acquisition press-release headlines and clickbait news headlines, and five large language models. We find that confidence thresholds reliably and substantially increase classification accuracy, albeit at the cost of more unresolved tasks, thus presenting practitioners an operational decision problem in optimizing the LLM-human workflow. Redundancy, although well established as effective at improving accuracy under human labeling, does not meaningfully improve accuracy while still increasing unresolved tasks. Combining a simple model of redundancy with the empirical evidence, we posit that this discrepancy stems from errors that are positively correlated within and across models—plausibly reflecting shared training data and architectures—violating the independence assumption that underlies the accuracy benefits of redundancy in human labeling. We translate these results into guidelines for a workflow design to create structured data from unstructured text: partition the corpus by self-reported confidence, accept model classifications for high-confidence cases, and route the remainder to human review, raising accuracy on the automated subset while concentrating human review where it adds the most value.
Human annotators routinely disagree on hate speech labels, especially when interpreting sarcasm, coded language, or platform-specific norms. This disagreement reflects genuine differences in community standards, not annotation error. Most detection systems discard it. Conventional models often treat this inter-annotator disagreement as label noise. We propose a Mixture-of-Experts architecture in which three QLoRA adapters, trained on microblogging (HateXplain), collaborative editing (Jigsaw), and extremist forum (CrimeBB) corpora, are blended via a learned gating mechanism atop a Llama 2 7B backbone. A control parameter adjusts detection sensitivity at inference time, enabling more conservative or more permissive predictions without retraining. Each prediction includes a structured rationale identifying targeted groups, toxic spans, and harm mechanisms. Training requires approximately 10 GPU-hours per adapter on a single NVIDIA RTX 4090 with 4-bit quantization. Across three benchmarks, the default Logit Blending configuration maintains Macro-F1 within 5 percentage points of a monolithic QLoRA baseline trained on pooled data, and discrete top-3 routing matches or exceeds it in accuracy on the extremist-forum and collaborative-editing benchmarks. Sensitivity tuning yields per-group accuracy gains of up to 6 percentage points, and blinded GPT-4 evaluation ranks blended rationales above the monolithic baseline in logical consistency, completeness, and fluency. The result is a framework for hate speech detection that yields a transparent and computationally efficient alternative to monolithic classification.
2Teaching
My teaching at UC Davis spans the Master of Science in Business Analytics (MSBA), MBA, and Online MBA programs.
Professor of the Year, MSBA Program, 2021, 2023, and 2026.
- Machine Learning & Artificial Intelligence
- MSBA, BAX-452. This course spans regression and regularization, tree-based methods, neural networks, language models, reinforcement learning, model selection, and evaluation.
- Data Design & Representation
- MSBA, BAX-422. This course develops methods for extracting, representing, and structuring unstructured data for analysis.
- AI & Business Innovation
- MBA, MGV-490A. This course connects regression and causal reasoning to predictive modeling, neural networks, and large language models.
- Agentic AI: From Models to Systems
- MBA, MGV-490B. This course builds agentic systems that coordinate models, process news, use external tools, and execute decisions through MCP.
- Data Wrangling
- MBA, MGT/B-435. This course turns heterogeneous web sources into reproducible, analysis-ready datasets through automated acquisition and processing.
- Business of the Future
- Wine Executive Program, University of California, Davis. This course examines how AI and data shape customer strategy, profitability, and managerial decision-making.
- Economic Research Experience for Undergraduates
- BA, Becker Friedman Institute, University of Chicago. This course introduces empirical economic research through programming, data construction, measurement, and analysis.
- New Tools for Acquisition / Analysis of Internet Data
- PhD, Harvard University and NBER. This course develops methods for turning internet information into usable data for empirical economics.
3Curriculum Vitae
I received my Ph.D. in Economics from the University of Chicago and hold degrees in mathematics and physics from Leipzig University. From 2015 to 2024, I held research appointments at Harvard University’s Center of Mathematical Sciences and Applications, first as a Postdoctoral Research Fellow and later as an Associate Research Fellow. I also serve as an Associate Editor for the Review of Economics and Statistics. Further details are in my curriculum vitae.
AAppendix: Google Scholar
Bibliographic records and citations to my work are available on Google Scholar.