Welcome to the Library (1400x200)

Archived Talks and Seminars

Reverse Chronological Order

Statistics and Data Science Seminar
Wednesday, September 23, 2026; 11:00am
Speaker Dr. Xialu Liu, Department of Management Information Systems, SDSU
Title Factor models with sparse loadings for high-dimensional time series
Abstract

High-dimensional data analysis using traditional models suffers from overparameterization. Two types of techniques are commonly used to reduce the number of parameters – regularization and dimension reduction. In this project, we combine them by imposing a sparse factor structure and propose a regularized estimator to further reduce the number of parameters in factor models. A challenge limiting the widespread application of factor models is that factors are hard to interpret, as both factors and the loading matrix are unobserved. To address this, we introduce a penalty term when estimating the loading matrix for a sparse estimate. As a result, each factor only drives a smaller subset of time series that exhibit the strongest correlation, improving the factor interpretability. The theoretical properties of the proposed estimator are investigated. The simulation results are presented to confirm that our algorithm performs well. We apply our method to Hawaii tourism data.

 

Statistics and Data Science Seminar
Wednesday, April 15, 2026; 11:00am
Speaker Dr. Tianchen Qian, Assistant Professor of Statistics, University Of California Irvine
Title Dynamic Causal Mediation Analysis for Micro-Randomized Trials
Abstract

Mobile health (mHealth) interventions aim to deliver real-time, personalized behavioral support based on the premise that prompts lead to short-term behavior changes, which in turn drive long-term health benefits. Verifying this premise empirically requires causal mediation analysis, but applying it to micro-randomized trials (MRTs), where individuals are randomized hundreds of times, poses two key challenges: (1) many treatment occasions (e.g., 210 decision points in the HeartSteps MRT with only 37 participants), making it difficult to define and model causal effects on an end-of-study outcome; and (2) many potential mediators, since a single treatment may influence the outcome through all future mediators, complicating the decomposition into interpretable pathways. We address these challenges by introducing natural direct and indirect excursion effects, which use marginal contrasts to define the causal effect of each treatment occasion on a distal outcome and focus mediation on the most immediate mediator following each treatment time. This captures the most scientifically meaningful pathways in mHealth (prompt → immediate behavioral target → long-term health goal) while improving the plausibility of causal identification assumptions. We derive efficient influence functions and propose a multiply robust estimator that accommodates flexible modeling and is doubly robust in the MRT context. Applied to the HeartSteps MRT, our analysis reveals that early direct effects of activity suggestions are large but decline rapidly, while the indirect effects, mediated through increased immediate walking, are more stable and persistent throughout the intervention.

 

Statistics and Data Science Seminar
Wednesday, March 4, 2026; 11:00am
Speaker Dr. Yuzhao Chen, Assistant Professor of Statistics, University Of California Riverside
Title Frontiers of Topology-guided Machine Learning Models for Spatio-temporal Data Learning
Abstract

Spatio-temporal event data play a central role in many modern scientific and societal applications, including urban mobility, public health, financial systems, and blockchain ecosystems. Such data are inherently relational, dynamically evolving, and non-Euclidean in nature, which poses fundamental challenges for classical statistical models and standard deep learning approaches. In this talk, I will introduce novel topology-guided machine learning frameworks that integrate spatio-temporal modeling, graph representation learning, and topological data analysis to address these challenges in a principled manner. First, I introduce a topology-enhanced diffusion framework for spatio-temporal point processes, which captures complex event dependencies through joint spatio-temporal graph construction and topological representations and leads to improved predictive performance and interpretability. Second, I discuss a multilayer topology-aware graph contrastive learning framework for fraud detection in time-evolving blockchain transaction networks, where persistent homology is employed to encode higher-order structural patterns beyond pairwise interactions. Together, these works demonstrate how topological structure can serve as a unifying analytical lens for learning from complex spatio-temporal data, and offer both practical performance gains and deeper statistical insight into the organization and dynamics of real-world complex systems.

 

Statistics and Data Science Seminar
Tuesday, February 10, 2026; 11:00am
Speaker Dr. Xiaowu Dai, Assistant Professor, Department of Statistics and Data Science, University Of California Los Angeles
Title Training-Free Multi-Agent Language Models
Abstract

Large Language Models (LLMs) have demonstrated strong generative capabilities but remain prone to inconsistencies and hallucinations. We introduce Peer Elicitation Games (PEG), a training-free, game-theoretic framework for aligning LLMs through a peer elicitation mechanism involving a generator and multiple discriminators instantiated from distinct base models. Discriminators interact in a peer evaluation setting, where rewards are computed using a determinant-based mutual information score that provably incentivizes truthful reporting without requiring ground-truth labels. We establish theoretical guarantees showing that each agent, via online learning, achieves sublinear regret in the sense their cumulative performance approaches that of the best fixed truthful strategy in hindsight. Moreover, we prove last-iterate convergence to a truthful Nash equilibrium, ensuring that the actual policies used by agents converge to stable and truthful behavior over time. Empirical evaluations across multiple benchmarks demonstrate significant improvements in factual accuracy. These results position PEG as a practical approach for eliciting truthful behavior from LLMs without supervision or fine-tuning.

 

Statistics and Data Science Seminar
Wednesday, February 4, 2026; 11:00am
Speaker Dr. Ting Fung Ma, Assistant Professor, Department of Statistics, University Of South Carolina
Title Hierarchical dependence modeling for the analysis of large insurance claims data
Abstract

Extreme weather events associated with climate change have caused significant damages. In particular, hail storms damage millions of properties in the U.S. and result in billion-dollar insured losses each year in the recent decade. To facilitate the insurance claims management operations in insurance companies, we construct a hierarchical dependence model, which accommodates the complex dependence within and between the outcomes of interests including the propensity of filing a claim, time to report a claim, and the claim amount. The storm-specific and property-specific characteristics are incorporated through marginal models, such as generalized linear models and survival analysis models. The dependence within the hail event is captured by spatial factor copula, while the dependence between different outcomes is captured by bivariate copula. For parameter estimation we develop a two-step procedure that first maximizes the marginal likelihood function and then maximizes the pairwise likelihood, which ensures computational feasibility for big data. We apply this modeling framework to analyze a large dataset involving hail storms in Colorado from 2011 to 2015 impacting hundreds of thousands of insured properties and demonstrate that the predictive performance can be improved by our proposed methodology.

 

Statistics and Data Science Seminar
Wednesday, November 12, 2025; 11:00am
Speaker Dr. Zhengyuan Zhu, Professor, Department of Statistics, Iowa State University
Title Optimal Integration of Data from Probability and Non-probability Sample
Abstract

Probability sampling has been the preferred approach for finite population inference for decades. In the big data era, nonprobability samples have become more prevalent due to their low data collection cost. Nonetheless, in the absence of a known inclusion mechanism, nonprobability samples may not accurately represent the target population without appropriate adjustments, and even a large non-probability sample may have a small effective sample size. To harness the strengths of both data sources, we develop a data integration method which enables a composite estimate using data from both probability and nonprobability samples when the variable of interest is observed in each. Our approach effectively mitigates informative selection and deterministic undercoverage in the nonprobability sample and addresses ignorable nonresponse in the probability sample. Simulation studies show that our optimal estimator is more efficient than those based solely on either sample type, and it outperforms several other alternative data integration estimators. We apply this method to analyze blood pressure data from US children and adolescents, integrating survey data from the National Health and Nutrition Examination Survey (NHANES) and administrative data from the Geisinger Health System. Variance estimation is provided using a replication method which accounts for the complex probability survey design of NHANES.

 

Statistics and Data Science Seminar
Wednesday, October 8, 2025; 11:00am
Speaker Cleridy Lennert-Cody, PhD, Senior Scientist at the Inter-American Tropical Tuna Commission (IATTC)
Title Adapting statistics to the real world: Estimation of target species catch for the tuna purse-seine fishery of the Eastern Pacific Ocean
Abstract

The tuna purse-seine fishery of the Eastern Pacific Ocean (EPO) currently produces over 20% of the Pacific Ocean catch of tropical tuna species (yellowfin, bigeye and skipjack tunas). The fishery is managed by an international fisheries commission, the Inter-American Tropical Tuna Commission (IATTC; www.iattc.org), whose headquarters are based in San Diego. To ensure that the catch of this fishery is sustainable, the scientific staff of the IATTC conduct population dynamics modeling for each of the three tuna species. Among other information, these population dynamics models require estimates of species catch amounts for the entire EPO purse-seine fleet. The species composition of the catch is estimated from sample data collected when vessels unload their fish in port. Improving upon the current sampling and estimation methodologies, which were developed 25 years ago for a fishery that has since evolved, serves as a reminder that the practical application of sampling in the real world can require flexibility in sampling design and innovative estimation methods. In this presentation we discuss the need for improvements to the current sampling protocol and estimation methods, a new sampling protocol proposed for implementation in 2026 and work in progress on new methods for species catch estimation.

 

Statistics and Data Science Seminar
Wednesday, September 17, 2025; 11:00am
Speaker Zhipeng Lou, Department of Mathematics, UC San Diego
Title Statistical inference of ranks
Abstract

Rank aggregation from pairwise and multiway comparisons has drawn considerable attention in recent years and has a variety of applications, ranging from recommendation systems to sports rankings to social choice. The existing literature on the ranking problem mainly concerns parameter estimation and algorithm implementation. However, there has been little investigation on the statistical inference theory of ranks. In this talk, I will start with a novel inference framework for ranks based on a modified Plackett-Luce model for multiway ranking with only the top choice observed. Then I will present a new methodology to construct simultaneous confidence intervals for the corresponding ranks through a sophisticated maximum pairwise difference statistic based on the MLE. Practically a valid Gaussian multiplier bootstrap procedure is developed to approximate the distribution of the proposed statistic. With the constructed simultaneous confidence intervals, we are able to study various inference problems on ranks such as testing whether an item of interest is among the top-K ranking. Our inference framework for the ranks can be widely applicable in many other ranking problems.

 

Statistics and Data Science Seminar
Wednesday, September 10, 2025; 11:00am
Speaker Joann Chen, Department of Computer Science, SDSU
Title Envisioning the Future of Digital Privacy
Abstract

In today's digital world, personal data represents both immense opportunity and significant risk. When used responsibly, it can drive breakthroughs in areas such as healthcare and finance, yet its misuse can result in severe privacy breaches. This talk will examine the evolving landscape of data privacy, highlighting Differential Privacy (DP) as a promising approach that offers strong, provable guarantees. It will further explore privacy risks in machine learning and the principles of DP-aware system design, with particular focus on the challenges of integrating DP into diverse real-world systems.

 

Statistics and Data Science Seminar
Wednesday, September 3, 2025; 11:00am
Speaker Xiyue Liao, Department of Mathematics and Statistics, SDSU
Title A Framework for Comprehensive Model and Variable Selection
Abstract

We propose a framework for choosing variables and relationships without assuming additivity or parametric forms. The relationships between the response and each of the continuous predictors are modelled with regression splines and assumed to be smooth and one of the following: increasing, decreasing, convex, concave, or a combination of monotonicity and convexity. The eight shapes include a wide range of popular parametric functions such as linear, quadratic, exponential, etc., and the set of choices is appropriate if the component functions "do not wiggle." An ordinal predictor can have its set of possible orderings, such as increasing, decreasing, tree or umbrella orderings, no ordering, or constant. Interactions between continuous predictors will be modelled as multi-dimensional warped-plane spline surfaces, where the same possibilities for shapes are considered. We propose combining stepwise selection methods with information criteria, LASSO-type ideas, and model selection using a genetic algorithm.