To help a (perhaps artificial) policy decider compare the performance of its (perhaps arti- ficial) social agents, a type of abstract evaluation machine is presented that could be used in various concrete operative situations. Designed on a sequence of progressively more challenging objectives on the chain O 0 , O 1 , ..., O l . . . ., O L , this machine, applied to any social agent A belonging to a given class, evaluates A’s performance taking into account the state x, provided that the gain scores ω l|x := V al |x (O l ) − V al |x (O l−1 ) which regulate the machine have been assigned. A criterion is proposed that interprets gain scores in a probabilistic context, by invoking an intuitive principle of “intrinsic worthiness”. The spe- cific question of interest concerns a structured, possibly automatable, approach to extract meaningful standardized gain scores ω l|x , from the “reference training-data set” associated to the choice of reference agent A ∗ (eg a “ good practice”), conditional also on assump- tions on modeling. An advanced (pseudo-Bayesian) approach is outlined in a Dirichlet- multinomial framework. A subspace M 0 of Dirichlet distributions (viewed as a geometric submanifold) is modeled by assuming that hidden processes, related to progress along the goal chain, are driven by hidden transition probabilities dependent on state x, governed by a set of hyper-hyperparameters. Applying an orthogonality principle, it is searched for that specific configuration of the hyperparameters set such that a certain (pseudo-Bayesian) vector-operator of updating, given the current training data set, is orthogonal to the sub- manifold M 0 ; thus tuning hyper-hyperparameters on the training-data set.
Tuning Social Indexes on Training-Data
D'EPIFANIO GIULIO
2026
Abstract
To help a (perhaps artificial) policy decider compare the performance of its (perhaps arti- ficial) social agents, a type of abstract evaluation machine is presented that could be used in various concrete operative situations. Designed on a sequence of progressively more challenging objectives on the chain O 0 , O 1 , ..., O l . . . ., O L , this machine, applied to any social agent A belonging to a given class, evaluates A’s performance taking into account the state x, provided that the gain scores ω l|x := V al |x (O l ) − V al |x (O l−1 ) which regulate the machine have been assigned. A criterion is proposed that interprets gain scores in a probabilistic context, by invoking an intuitive principle of “intrinsic worthiness”. The spe- cific question of interest concerns a structured, possibly automatable, approach to extract meaningful standardized gain scores ω l|x , from the “reference training-data set” associated to the choice of reference agent A ∗ (eg a “ good practice”), conditional also on assump- tions on modeling. An advanced (pseudo-Bayesian) approach is outlined in a Dirichlet- multinomial framework. A subspace M 0 of Dirichlet distributions (viewed as a geometric submanifold) is modeled by assuming that hidden processes, related to progress along the goal chain, are driven by hidden transition probabilities dependent on state x, governed by a set of hyper-hyperparameters. Applying an orthogonality principle, it is searched for that specific configuration of the hyperparameters set such that a certain (pseudo-Bayesian) vector-operator of updating, given the current training data set, is orthogonal to the sub- manifold M 0 ; thus tuning hyper-hyperparameters on the training-data set.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


