Changelog
Source:NEWS.md
tseLCA 2.0.0
Step-wise interface
- New
tse_lca()fits the Step-1 measurement model from a formula,cbind(Y1, Y2, ...) ~ 1, with acontrol = tse_control()argument. With several numbers of classes (e.g.nclass = 1:6) it returns atseLCA_selectclass-enumeration table: log-likelihood, parameters, AIC, BIC, SABIC, entropy R^2, and smallest class. The table hasprint(),plot(),as.data.frame(),best_model(), and[[methods. The one-class (independence) model is fitted in closed form. - New
tse_classify()(Step 2) assigns observations to classes with a fixed measurement model (assignment = "modal"or"proportional"). It reports the classification-error probabilities P(W = s | X = t) and entropy R^2. Withnewdatait classifies another sample, replacing thestep1 =route for using a measurement model from one sample on another.tse_lca()models now keep theirdatafor this purpose. - New Step-3 functions:
-
tse_covariate()fits the multinomial logit of class membership on a covariate formula (factors, interactions, transformations). Arguments:method = "ML","BCH", or"none"(the uncorrected three-step estimator, for comparison);se = "corrected"or"robust";ref(reference class);start. Methods:predict()(class probabilities given covariates),relevel(), andanova()(Wald tests per term). -
tse_distal()fits a distal outcome (Zo ~ 1, any supportedfamily, as a string or family object). Given atse_covariate()model it fits the combined model. -
tse_twostep()gives two-step estimates (Bakk & Kuha 2018). Withse = TRUEthey come with multilevLCA’s corrected standard errors.
-
- New
tseLCA()fits a whole model in one call:cbind(indicators) ~ covariates | distal outcome(Formula package). It chains the step-wise functions. Accessorsmeasurement(),classification(),covariate(), anddistal()extract the components of any fitted model. -
tse_covariate()andtse_distal()acceptdata: the classified data (same rows) with additional columns, e.g. variables created after classification. -
omnibus_test()returns a standardhtestobject. - The vignette, README, package help page, and pkgdown reference index are rewritten for the step-wise interface. The vignette ends with a table mapping
three_step()arguments to the new functions. -
predict()for measurement models gives posterior class probabilities (or modal classes) for new data, coded with the model’s stored categories.fitted(),formula(), andupdate()also work. -
item_probs(se = TRUE)andclass_sizes(se = TRUE)also return standard errors (delta method from the measurement model’s variance). - Fitted objects no longer keep multilevLCA’s unused observation-level score matrix, which made up most of their size (about 80% smaller).
-
startintse_lca()andtseLCA()acceptsitem_probs()of a fitted model directly, e.g. to refit a chosen model with covariates. - New
as_tse_lca()builds a measurement model from given class sizes and item-response probabilities (e.g. estimates reported elsewhere, or saved from an earlier fit) without re-estimating it. The result can be passed totse_classify()like anytse_lca()model. - The replication and simulation scripts of the accompanying manuscript are no longer installed with the package; they are part of the manuscript’s replication materials.
Classes and methods
- Fitted objects now share a class hierarchy.
tseLCAis the parent class of every fit;tseLCA_structuralis the parent oftseLCA_covariate,tseLCA_distal, andtseLCA_both. Methods are defined once on the parent classes, not separately for each subclass. - New methods on all fits:
logLik()(withdfandnobs, soAIC()andBIC()work),nobs(),posterior(),classes(),class_sizes(), anditem_probs(). -
summary()now returns asummary.tseLCA_structuralorsummary.tseLCA_measurementobject. Its tables print withprintCoefmat()and are extracted withcoef(summary(fit)). -
plot()now works on distal-outcome fits. These objects previously did not store the measurement model, so the plot failed.
Data input
- Indicators can be factors, logicals, character, or numeric codes in any coding (e.g. 1..K). They are recoded to 0..K-1 internally, and their categories are stored with the measurement model (
$measurement_model$Y.levels). Previously they had to be integers starting at 0. Codes such as 1..K were accepted but silently misread: the highest category was never matched. - A measurement model reused through
step1applies its own categories to new data. The new data may omit categories, and values outside them are an error. - Covariate designs are built with
model.frame()/model.matrix(), so factor covariates are dummy coded and unused factor levels are dropped. The intercept is named"(Intercept)"(was"Intercept"). - Distal outcomes are validated for their family. Binomial outcomes may also be logical or two-category factors/characters.
- New
tse_control()collects the numerical estimation settings.
Bug fixes
tse_lca(missing = "fiml")with random starts (n_init) failed with a cryptic error from multilevLCA (“sort_index(): detected NaN”) when indicator values were missing: multilevLCA cannot start from a given classification on incomplete data. It now warns and uses multilevLCA’s default initialization; a user-suppliedstartgives a clear error.Corrected (ML) standard errors: the Jacobian of the classification-error matrix with respect to the Step-1 parameters, which carries the Step-1 uncertainty into Step 3, had its item-parameter columns ordered item by item, while the Step-1 variance is ordered class by class; with modal assignment it also differentiated the assignment as if it were the posterior probabilities. Both are fixed (checked against numerical derivatives). Point estimates, robust, and BCH standard errors are unchanged; corrected covariate standard errors typically increase by a few percent, and corrected standard errors no longer depend on the choice of reference class.
Structural models with a reference class other than the first (
ref =):posterior(),classes(),item_probs(),class_sizes(), andplot()reported the classes in the rebased order under the original labels, and in combined covariate + distal models the distal parameters of class t belonged to another class. All now use the measurement model’s class order.Step-1 variance of polytomous indicators when the first (reference) category of an item has a boundary probability in some class: the free parameters of that item are log-ratios against this category, and their common shift, log P(Y = first | class), is not informed by the data. It is now treated as fixed, like other boundary parameters. Previously the variance of these log-ratios was huge (the information matrix was nearly singular) or all
NA(singular), which made the corrected standard errors of the Step-3 modelsNA.If the Step-1 information matrix is still singular, Step-3 models warn and report robust standard errors (was
NA).tse_twostep(se = TRUE)with a reference class other than the first returned the class-1-reference estimates and variance under the new class labels. They are now transformed to the requested reference class. It also no longer warns, wrongly, that multilevLCA’s measurement model differs.For measurement-only fits,
posterior()/$posteriorsandclasses()/$classificationswere not in data-row order. They were read from multilevLCA’sfit0$mU, which is sorted by response pattern. They are now computed from the Step-1 data in data-row order. Code that used$classificationsfrom a measurement model, e.g. asstartval, received mismatched rows in 1.1.x.Polytomous indicators were decoded from
fit0$mUas category 0 whenever the measurement model was fitted with listwise deletion (incomplete = FALSE). The Step-1 information matrix was then singular, sovcov()of a polytomous measurement model and the Step-1-corrected Step-3 variance (use.simple.cov = FALSE) were allNA. The measurement model now stores its Step-1 sample in data-row order and uses it for all of these computations. The fallback decoder for objects without stored data handles both of multilevLCA’s codings.Parameter counts behind the reported AIC/BIC were wrong. Covariate models counted one-hot indicator columns, not free item parameters (e.g. 40, not 22, for six binary items, three classes, one covariate). Combined covariate + distal models counted class sizes, not the covariate coefficients. Reported
AIC/BICfor these models change; estimates and standard errors do not.Gaussian distal outcomes: the within-class variance was fixed at sigma2 = 1 in the Step-3 likelihood. Three-step ML now estimates a common sigma2 jointly with the class means, as in Bakk, Tekle & Vermunt (2013). Unlike ordinary regression, sigma2 does not factor out of the mean estimates here: it enters the posterior weights P(X = t | W, Zo). With the variance fixed at 1, ML class means were biased whenever the outcome’s variance differed from 1 (e.g. roughly 3.5x too far apart for a residual SD of 3) and were not equivariant to rescaling the outcome. The Hessian, score, and Step-1/Step-2 uncertainty propagation now include sigma2. It is reported as
$sigma2(estimate and, under ML, standard error), and counted in the log-likelihood’s degrees of freedom. BCH class means were unaffected; their log-likelihood now uses the estimated variance.Three-step ML for distal outcomes with proportional assignment (
use.modal.assignment = FALSE) maximized the wrong likelihood. It used log sum_t P(t) f(z|t) sum_s w_s P(W=s|t), with the assignment weights inside the log. The likelihood of Bakk, Tekle & Vermunt (2013), for the expanded data file weighted by the posterior assignment probabilities, is sum_s w_s log sum_t P(t) f(z|t) P(W=s|t). The covariate model already used this form. The E-step, score, Hessian/Jacobian, and Step-1/Step-2 uncertainty propagation now use the correct likelihood for all distal families. Proportional-assignment ML distal estimates were biased away from zero (e.g. class means of -1.07 and 1.06, not -1 and 1, at mid separation). Modal-assignment fits, whose single assignment makes the two forms identical, BCH, and covariate models are unchanged.Covariate models with
include.intercept = FALSEfailed with a “length of ‘dimnames’” error. The coefficient rows were always labeled with anInterceptrow.In combined covariate + distal models, rows with a missing covariate but an observed distal outcome were kept in the distal model with
NAclass priors. The distal model now uses rows with complete covariates.
Deprecated
-
three_step()is deprecated in favor oftseLCA()and the step-wise functions. It keeps working, with the same estimates, and warns once per session. Its help page maps each argument to the 2.0 interface. -
lca_step1(),lca_step1_startval(),fitZ_from_fit0(), andfitZ_from_multiLCA()are deprecated in favor oftse_lca()andtse_twostep(). They warn once per session when called directly. -
options(tseLCA.warn.deprecated = FALSE)silences these warnings.
Breaking changes
-
coef()returns a named vector whose names matchvcov(), soconfint()works.coef(fit, matrix = TRUE)gives the previous Q x (T-1) (covariate) or T x C (multinomial distal) layout. -
coef()on a measurement model returns the log-ratio parameters thatvcov()describes. The previous list is now available asclass_sizes()anditem_probs(). - Covariate coefficient names use
"(Intercept)"(was"Intercept"), e.g."(Intercept):C2". - The
whichargument ofcoef()andvcov()is replaced bycomponent("covariate"or"distal") andstep("two_step"). Passingwhichgives an error, so old code cannot silently return a different quantity. -
vcov()on atseLCA_bothobject returns one matrix. The covariate and distal blocks are on the diagonal; the cross-covariances are not computed and areNA.
tseLCA 1.1.1
CRAN release: 2026-09-23
- Corrected Jay Goodliffe’s role in
Authors@Rfrom contributor ("ctb") to author and copyright holder (c("aut", "cph")), reflecting his contribution to the package. No code changes.
tseLCA 1.1.0
CRAN release: 2026-09-20
Externally supplied Step-1 starting values
- Added
lca_step1_startval(), a wrapper aroundmultilevLCA::multiLCA()that injects a user-supplied Step-1 starting value and setskmea = FALSE, bypassingmultilevLCA’s default deterministic k-means-on-principal- components initialization. This addresses feedback that the default initialization can consistently converge to a local optimum of the Step-1 log-likelihood on some datasets.startvalaccepts either:- an integer classification vector (
1..n_classes, one entry per row ofdata), e.g. the modal class from an external solver run with many random starts (‘StepMix’, ‘poLCA’); or - a numeric matrix of conditional item-response probabilities
P(Y_h = k | X = t), from which a classification is derived internally (naive-Bayes argmax under a flat class prior). This is the natural format for an externally estimated Step-1 solution that isn’t tied to the current sample, e.g.poLCA’sprobsoutput or a published item-response table. Either way, users can pass their external Step-1 solution straight through instead of hand-assembling atseLCAmeasurement object.
- an integer classification vector (
- Added a matching
startvalargument tolca_step1()andfitZ_from_multiLCA(), so every place a measurement model is fit from raw data can use an external starting value instead ofmultilevLCA’s k-means initialization. Supplyingstartvalskips theiter.measurement/R2.thresholdrandom-restart logic, since restarting from fresh k-means seeds would defeat the purpose of a user-vetted start. - Added a
startvalargument tothree_step(), forwarded tolca_step1()(and tofitZ_from_multiLCA()whenget.twostep.vcov = TRUE) so the measurement model can be fit from an external starting value in a single call.
Multiple random-start Step-1 initialization (n_init)
- Added an
n_initargument tolca_step1(),fitZ_from_multiLCA(), andthree_step(): the unconditional multi-random-start analog ofn_initinStepMixornrepinpoLCA. When supplied, the measurement model is fitn_inittimes from independent uniform-random classifications (kmea = FALSE, notmultilevLCA’s deterministic k-means-on-PCA path), and the highest-log-likelihood fit is kept. This is a separate mechanism from the existingiter.measurement/R2.thresholdrestart logic, which rerunsmultilevLCA’s own k-means initialization and only when entropy R^2 is low;n_initrestarts always run and never use k-means. -
step1,startval, andn_initare mutually exclusive ways of controlling Step 1 inthree_step()(andstartval/n_initinlca_step1()andfitZ_from_multiLCA()); supplying more than one errors.
Bug fixes
- Fixed the orientation of the BCH weight matrix used throughout
use.bch = TRUEestimation (both covariate and distal-outcome models).pwx[s, t] = P(W = s | X = t)is column-stochastic; the correct BCH weight matrix isw.is %*% t(pwx)^-1(Mplus Web Note 21: the row ofH^-1for each case’s most likely class, whereH = t(pwx)is row-stochastic), notw.is %*% pwx^-1, which the code had been computing. The two orientations only agree whenpwxis symmetric, so the bug was largely invisible in well-separated, balanced test cases; with genuine classification-error asymmetry it biased BCH point estimates and could produce the “negative column sums” error the package warns about (users were advised to fall back touse.bch = FALSE). With the corrected orientation, every row of the weight matrix sums to 1 and the class totals it implies exactly match the posterior class sizes. The weight-matrix computation is now consolidated into a single internal helper (bch_weight_matrix()) used by all four call sites that previously duplicated it, with a regression test assertingrowSums(w.it) == 1under a deliberately asymmetric classification-error matrix.
Multinomial distal outcomes and an omnibus class-equality test
- Added
family = "multinomial"tothree_step()’s distal-outcome estimation, for a nominal categorical outcome with 2 or more categories (Zo.namemay be a factor, character, or integer column). Estimates a saturated model – theT x Cmatrix of class-conditional category probabilitiespi_hat[t, c] = P(Zo = c | X = t)– with the closed-form weighted-proportion estimatorpi_hat[t, c] = sum_i w_it * 1(y_i = c) / sum_i w_it, for bothuse.bch = TRUE(BCH weights) anduse.bch = FALSE(ML, through EM with the same closed-form M-step and responsibility-weighted E-step). This replaces the previous workaround of fitting onefamily = "binomial"model per category and renormalizing the resulting probabilities by hand, which wasn’t constrained to the simplex before renormalizing and had no joint covariance across categories.-
coef()returns theT x Cprobability matrix (rows sum to 1) instead of a length-Tvector;vcov()returns its(T*C) x (T*C)sandwich covariance, which is necessarily rank-deficient (each class’s row sums to 1). - For
use.bch = FALSE,use.simple.cov = FALSE(the default) fully propagates both Step-1 measurement uncertainty and, whenZp.namesis also supplied, Step-3 covariate uncertainty into the SEs, matching the existing gaussian/poisson/binomial ML paths exactly (the T x C generalization of each chain-rule term –C1_matfor Step 1,C_matfor Step 2 – only required expanding the “unit score”g_itacross categories, since neither term’s derivation otherwise depends on the outcome’s dimensionality). The ML bread (multinomial_ml_jacobian()) is a closed-form Jacobian of the estimating equation, generalizingml_hessian_distal()’s “observed = complete - missing information” correction to the T x C case. Both the bread and the full propagation (including the Step-2 term) were cross-validated against independent, from-scratch numerical differentiation (explicit per-case loops sharing no code with the package,optim()from multiple starting points for the point estimate) and matched to machine precision.
-
- Added
omnibus_test(), a generalized Wald test ofH0: theta_1 = theta_2 = ... = theta_T(the distal outcome’s class-t parameter vector is the same for every class) for anytseLCA_distalortseLCA_bothobject, regardless of family. Uses a Moore-Penrose pseudo-inverse of the contrast covariance so it remains valid when that covariance is singular, as it always is forfamily = "multinomial"; the resulting degrees of freedom recover the textbook(T-1)*(C-1)for aT x Cchi-squared test of homogeneity in that case, andT-1for the scalar-parameter families. - Unlike
family = "binomial", whosecoef()/vcov()are on the logit scale,family = "multinomial"reportscoef()/vcov()directly on the probability scale –Std.Erroris directly interpretable without a delta-method back-transform, but a symmetric intervalEstimate +/- 1.96*Std.Errorcan fall outside the unit interval for a probability near a boundary, the same known limitation as a naive Wald interval for a sample proportion, and the per-cellz.value/p.value(testing each probability against 0) are rarely the question of interest.print()/summary()now print a one-line reminder of this after a multinomial distal-outcome table, pointing toomnibus_test()for the intended, boundary-safe test of whether the distribution differs across classes.
tseLCA 1.0.0
- Initial submission to CRAN.
Core Estimation Framework
- Implemented BCH and ML bias-adjusted three-step estimators for latent class analysis (LCA).
- Added support for structural models containing covariates (), distal outcomes (), and combined models (estimating the relationship between and the latent class first, followed by the distal outcome adjusting for covariate-adjusted posteriors).
- Implemented analytic sandwich variance estimation to correctly propagate measurement uncertainty from the first-step LCA through classification-error correction in the final step.
- Added a robust standard error option (
use.simple.cov = TRUE) that bypasses the measurement-uncertainty correction for faster computation in large, well-separated samples.
Measurement Model (Step 1) Integration
- Integrated with the ‘multilevLCA’ package for efficient Step-1 measurement model estimation.
- Added support for polytomous indicator items (0-based integer coding).
- Implemented Full Information Maximum Likelihood (FIML) to handle missing data in the measurement model with the
incomplete = TRUEargument (using a two-pass row-filtering strategy). - Added the ability to pass a pre-fitted measurement model (with the
step1argument) to reuse across multiple structural models or apply to different sample subsets. - Implemented automated random restarts for the measurement model triggered when entropy falls below a user-specified threshold.
Algorithmic Flexibility & Structural Models
- Added support for both modal and proportional (soft) posterior class assignment (
use.modal.assignment). - Integrated Gaussian, Poisson, and binomial families for distal outcome estimation.
- Added the
rebaseargument to allow users to easily change the reference latent class for the multinomial logit parameterization while maintaining invariant log-likelihoods. - Implemented two-step EM estimation (
fitZ_from_fit0()) to generate stable starting values for the three-step structural model.
Utilities and Methods
- Included standard S3 methods for
tseLCAobjects:summary(),coef(),vcov(), andplot()(which delegates to ‘multilevLCA’ for item-profile visualization). - Built a data-generating process (
generate_data()) that replicates the Bakk & Kuha (2018) simulation study design for both covariates and distal outcomes under varying separation conditions.
tseLCA 1.0.1
CRAN resubmission
- Removed single quotes around acronyms in DESCRIPTION; added explanations of BCH, ML, and LCA.
- Replaced
T/FwithTRUE/FALSEthroughout internal codebase - Added
\valuetags to all exported functions missing them, includingbk2018_params. -
inst/examples: examples now write totempdir()instead of the home filespace. -
inst/examples: commented outrm(list = ls())calls. -
inst/examples: commented outinstall.packages()calls.