Building probability models from statistics by combining distributions

Hiroshi Yamashita, Hideyuki Suzuki, and Kazuyuki Aihara, “Entropic herding,” Statistics and Computing 33, 31 (2023).

Sometimes we know averages or correlations without having the individual observations. What probability distribution can we construct from those statistics? Herding generates a sequence of points whose feature averages approach specified targets. Entropic herding extends this idea by generating a sequence of tractable distributions, such as Gaussians. It updates feature weights according to discrepancies in the statistics and uses those weights to select the next distribution. Combining the components produces a mixture that can represent multiple modes and dependencies, even when each component is simple.

The central contribution is a derivation from an objective combining moment error with entropy, clarifying herding’s connection to the maximum entropy principle. The paper characterizes the objective’s optimum and recovers conventional herding by restricting the candidates to point distributions and choosing the update rule appropriately. It also explains how diversity among the generated components can increase the mixture’s entropy, complementing the explicit entropy term in each optimization step.

With suitable components, the resulting model supports density evaluation, likelihood-based validation, and independent sample generation. Numerical examples include a bimodal distribution, a Boltzmann machine, and wine data used for classification and conditional inference of missing values. Finite penalty weights and approximate optimization leave discrepancies from the ideal maximum-entropy distribution; the paper does not establish a general convergence guarantee. Its contribution is to extend moment-matching herding into a method for constructing usable probability models.

Target statistics guide sequential distributions and feature-weight updates to build a mixture model.