Research

Papers and Projects

Manuscripts Under Review

Google Scholar

Under review at Operations Research

Stochastic Shortest Path Interdiction with Path-Only Feedback

We study a sequential stochastic shortest-path interdiction problem under path-only feedback, where the interdictor observes the evader's chosen route but not the realized arc costs that generated it. The paper characterizes the resulting identifiability problem, derives an instance-dependent logarithmic regret lower bound, and develops Identifying-Exploration Greedy policies that attain logarithmic regret on identifiable instances. Numerical experiments include a border-infiltration case study on the Arizona-Mexico network.

Keywords
  • Network interdiction
  • Sequential learning
  • Multi-armed bandits
  • Partial monitoring
Read on SSRN Download PDF
Arizona–Mexico infiltration network with numbered nodes and the Gulf of California
Arizona–Mexico computational network; the Gulf of California appears at lower left.

Working Papers

Observed inventory snapshots surround hidden, sampled customer paths.

On the Estimation of a Markov Chain Choice Model from Period-Level Inventory Observations

Abstract

We study the estimation of Markov chain choice models (MCCMs) from period-level inventory observations, a common data regime in which the analyst observes only the initial inventory, final inventory, and number of arrivals for each selling period—not individual purchases, the assortments faced by customers, or the order in which sales and no-purchases occurred. We formulate the observed-data likelihood by introducing the latent period-level purchase path and extend the expectation-maximization (EM) framework for the MCCM to this setting. The central computational difficulty is that likelihood evaluation and the E-step both require averaging over the same combinatorially large set of feasible paths, but under different distributions: uniform sampling suffices for likelihood evaluation, while the E-step must target the posterior distribution over paths induced by the current parameter estimates. We develop two posterior path samplers: a Metropolis–Hastings scheme with importance-sampling reuse across EM iterations, and a sequential Monte Carlo (SMC) sampler that constructs paths forward using look-ahead proposals guided by the observed final inventory. We establish an identifiability result and an exact-EM convergence guarantee for the underlying deterministic algorithm. Numerical experiments show that uniform path sampling is adequate when stockouts are rare but degrades sharply once demand exceeds supply, because it misweights the arrival orders that determine when stockouts occur; posterior sampling corrects this bias, yielding substantial improvements in inventory, stockout, likelihood, and parameter-recovery metrics, both under a well-specified MCCM and under ranking-based misspecification of the demand process. Between the two posterior samplers, SMC typically offers better runtime performance without sacrificing accuracy.

Keywords
  • Markov chain choice models
  • Inventory data
  • Expectation-maximization
  • Posterior path sampling
  • Sequential Monte Carlo
Independent product draws meet in one shared assortment-level choice.

Avoiding Exponential Dependence in Thompson Sampling for Multinomial Logit Bandits

Abstract

Dynamic assortment learning under multinomial logit (MNL) demand couples revenue and information through a shared choice denominator. Product-versus-outside comparisons support per-customer updates, but existing worst-case analyses of MNL Thompson sampling correlate the product draws, because the probability that several independently sampled coordinates are optimistic at once can decay exponentially in the display capacity K. We show that such correlation is not necessary. In linear-fractional form, assortment optimism is a single signed linear inequality in the attractions, so favorable draws offset unfavorable ones. Combining this assortment-level optimism with martingale concentration for per-product clocks that permit revision after every customer, we obtain gap-free frequentist regret Õ(K√(NT)) for independent, variance-inflated draws, with no exponential capacity penalty; the same draw attains the same order under the original epoch clock. Retaining instead the complete item-level MNL likelihood, which these clocks discard, we construct a Gibbs sampler for the joint posterior and prove that ideal Thompson sampling from it has Bayesian regret Õ(√(NT)), uniformly in K. To our knowledge, these are the first such guarantees for product-wise independent draws and for the complete MNL posterior. A computational study compares the filtered, mean-field, and complete-posterior policies, separating posterior dependence from correlation imposed across exploration draws.

Keywords
  • Multinomial logit bandits
  • Thompson sampling
  • Dynamic assortment learning
  • Bayesian regret
  • Gibbs sampling
Two vaccination cohorts converge on the same linked hospital endpoint.

Effectiveness of a Second Dose of Nirsevimab against RSV Hospitalisation: A Two-Cohort Natural Experiment in Chile

Abstract

We evaluate the effectiveness of a second dose of nirsevimab among infants entering a second respiratory syncytial virus season. Using a two-cohort natural experiment in Chile, the study links national health records to compare second-season RSV hospitalisation among children eligible for one versus two doses.

Keywords
  • Nirsevimab
  • Respiratory syncytial virus
  • Natural experiment
  • Second-season protection
  • Linked health records

M.Sc. thesis · In progress

Thompson Sampling for Multinomial Logit Bandits: Learning in Monopolistic and Competitive Markets

Read Abstract

When customer preferences are unknown and a seller must decide in each period which products to offer, every assortment decision plays a dual role: generating revenue and revealing information about demand. This interaction between decision-making and learning gives rise to a bandit problem. We consider demand governed by a multinomial logit (MNL) model, whose shared denominator makes both purchase probabilities and the information conveyed by each choice depend on the entire assortment. Under competition, this assortment also includes rival products; consequently, other sellers’ decisions affect both a seller’s revenue and what it can learn when it does not make a sale.

This thesis studies Thompson sampling policies for this problem with either a single seller or multiple sellers learning in a decentralized manner. All policies sample a plausible vector of attraction parameters and act as if it were the true vector. They differ in how they represent uncertainty: filtering the history into product-versus-outside-option comparisons to obtain independent Beta-prime updates; preserving posterior dependence through a Gamma–Exponential representation; or approximating the shared denominator to maintain a separate state for each product.

For a single seller with a catalog of N products, capacity to offer at most K of them, and a horizon of T periods, two of the proposed policies admit regret bounds. The policy that samples attractions independently achieves a bound of order Õ(K√(NT)), whereas the policy that samples exactly from the joint posterior attains Bayesian regret of order Õ(√(NT)). Under competition, each seller combines its own draw with the potential function of the MNL game. With full observation, the expected number of periods outside the pure-strategy Nash equilibrium selected by the potential is o(T); with partial observation, the expected number of periods outside the equilibrium set is also o(T), provided that decisions are identifiable. Taken together, these results show that learning demand and stabilizing a competitive market depend on the information observed, the extent to which MNL dependence is preserved, and the computational burden one is willing to bear.

Conference Presentations

On the Estimation of a Markov Choice Model from Partial Inventory Data

2026 INFORMS Annual Meeting · San Francisco, USA

Accepted · November 2026

On the Estimation of a Markov Choice Model from Partial Inventory Data

XXIII Latin-Iberoamerican Conference on Operations Research · Bogotá, Colombia

Accepted · November 2026

Economic Evaluation of a Universal Screening Policy for Congenital Cytomegalovirus in Chile

44th Annual Meeting of ESPID · Bologna, Italy

Poster · June 2026

On the Estimation of a Markov Choice Model from Partial Inventory Data

XV Chilean Conference on Operations Research · Coquimbo, Chile

December 2025