Users comparing fit statistics across versions of
OpenMx, umx, and lavaan may
notice differences in reported \(\chi^2\), CFI, TLI, and RMSEA values for
Weighted Least Squares (WLS / DWLS) models.
These differences do not reflect arbitrary code changes. They stem from distinct statistical formulations designed for different estimation contexts: 1. Raw WLS Discrepancy: The unadjusted quadratic form fit function. 2. Browne (1984) Residual \(\chi^2\): Discrepancy evaluated directly on the asymptotic covariance of sample moments. 3. Satorra–Bentler (1994, 2010) Scaled \(\chi^2\): Mean-and-variance scaling applied to WLS models for omnibus hypothesis testing. 4. Savalei (2021) CatML Corrections: Evaluating WLS parameter estimates inside a Maximum Likelihood correlation scaffold to obtain valid, non-inflated CFI, TLI, and RMSEA fit indices for ordinal data.
This guide explains the mathematical rationale, history, exact
implementation, comparative logic between umx and
lavaan, and practical manuscript reporting guidance.
Under Maximum Likelihood (ML), variables are assumed continuous and multivariate normal. When data are ordinal (binary or Likert), structural equation models estimate polychoric correlations and thresholds for underlying continuous latent responses \(x^*\).
In WLS/DWLS estimation: * The sample polychoric correlation matrix \(R\) has exact \(1.0\)s on the diagonal. * Binary indicators fix their residual variance to \(1.0\). Consequently, the model-implied total variance of \(x^*\) is \(\text{Var}(x^*) = 1.0 + \text{explained variance} > 1.0\).
This metric difference between observed correlation \(R\) (\(\text{diag}=1\)) and model-implied response covariance \(\hat{\Sigma}\) (\(\text{diag} > 1\)) creates distinct challenges for fit evaluation.
The raw WLS discrepancy is given by: \[F_{\text{WLS}} = (s - \hat{\sigma})^T W (s - \hat{\sigma})\] where \(s\) is the vector of sample moments (polychoric correlations and thresholds), \(\hat{\sigma}\) is the model-implied vector, and \(W\) is the weight matrix (full or diagonal).
OpenMx exposes an internal slot output$chi based on
Browne’s WLS residual weighted least squares: \[\chi^2_{\text{Browne}} = e^T \hat{V}^{-1}
e\] This statistic evaluates discrepancy directly in the space of
the asymptotic covariance matrix \(\hat{V}\).
The standard Satorra–Bentler scaled statistic adjusts the raw WLS objective by a scaling factor \(c\): \[\chi^2_{\text{SB}} = \frac{F_{\text{WLS}}}{c}\] where \(c = \text{tr}(U \Gamma) / r\), with \(U\) representing the residual weight matrix, \(\Gamma\) the asymptotic covariance of sample moments, and \(r\) the model degrees of freedom.
umxSummary() reports \(\chi^2_{\text{SB}}\) as the omnibus test
statistic for overall model fit when WLS Jacobians are available.
For decades, software packages calculated CFI, TLI, and RMSEA directly from WLS \(\chi^2_{\text{SB}}\). However, Victoria Savalei (2021) demonstrated that WLS \(\chi^2\) statistics under model misspecification do not follow a non-central \(\chi^2\) distribution in the same way ML \(\chi^2\) statistics do. As a result: * Conventional Hu & Bentler (1999) cutoffs (e.g., CFI \(> 0.95\), RMSEA \(< 0.06\)) do not apply to unadjusted WLS fit indices. * Standard WLS fit indices are severely distorted when binary and multi-category items are combined.
To restore valid, interpretable fit indices where conventional cutoffs apply, Savalei proposed evaluating the converged WLS parameter estimates \(\hat{\theta}_{\text{WLS}}\) inside a Maximum Likelihood correlation fit function (\(catML\)): \[XX_3 = n \cdot F_{\text{ML}}(R, R_\theta)\]
Because binary indicators produce model-implied response covariances with diagonal elements \(> 1.0\), evaluating \(F_{\text{ML}}(R, \hat{\Sigma})\) directly creates an artificial metric penalty.
umx solves this by standardizing the manifest response
variables before catML evaluation: 1. Standard deviations \(s_i = \sqrt{\hat{\Sigma}_{ii}}\) are
calculated for each manifest variable. 2. A diagonal scaling matrix
\(D_s\) is constructed (\(s_i\) for manifests, \(1.0\) for latents). 3. The RAM matrices are
rescaled: \[A^* = D_s^{-1} A D_s\]
\[S^* = D_s^{-1} S D_s^{-1}\]
This matrix transformation proves mathematically that: \[\text{diag}(\hat{\Sigma}^*) = 1.0\] The rescaled model yields a true correlation matrix expectation \(R_\theta\), allowing \(XX_3\) to be evaluated on the proper metric.
Savalei’s robust scaling factor \(\hat{c}_3\) is then applied to \(XX_3\) to compute robust CFI, TLI, and RMSEA values that honor standard Hu & Bentler cutoffs.
umx (Savalei catML)
vs. lavaan (WLSMV)When fitting the exact same model (identical factor structure,
identical degrees of freedom \(\text{df} =
24\), and identical sample data), users may notice differences in
reported \(\chi^2\) and fit indices
between lavaan and umx.
lavaan WLSMV
(chisq.scaled): Takes the unscaled DWLS
discrepancy \(T_{\text{DWLS}} = 42.99\)
and divides by a Mean-and-Variance scaling factor \(c_{\text{WLSMV}} = 0.736 < 1.0\): \[\text{chisq.scaled}_{\text{lavaan}} =
\frac{42.99}{0.736} = \mathbf{62.23}\] Because \(c < 1.0\), dividing \(42.99 / 0.736\) inflates
lavaan’s reported \(\chi^2\).
umx Savalei (2021) catML (\(\chi^2_{\text{Savalei}}\)):
Evaluates catML discrepancy \(XX_3 =
77.16\) at the WLS estimates and applies the Savalei (2021)
correction factor \(c_3 = 1.815\):
\[\chi^2_{\text{Savalei}} =
\frac{77.16}{1.815} = \mathbf{42.52}\]
Key Insight: Notice that \(42.52\) (
umx) and \(42.99\) (lavaanunscaled \(T_{\text{DWLS}}\)) represent the exact same underlying model discrepancy (\(\chi^2 \approx 42.5 - 43.0\) on \(\text{df} = 24\)).
Why does lavaan report CFI \(=
0.961\) while umx reports CFI \(= 0.758\) on datasets with mixed
categorical levels (e.g., HSwls, where \(x_1-x_3\) are binary, \(x_4-x_6\) have 3 levels, and \(x_7-x_9\) have 4 levels)?
Baseline \(\chi^2\)
Inflation in Standard WLSMV: In datasets with mixed ordinal
categories, standard DWLS weight matrices (\(\mathbf{W}\)) weight binary items
differently from 3+ level items. Under standard WLSMV in
lavaan, this causes the independence (null) model \(\chi^2\) to explode (\(\text{chisq.scaled}_{\text{baseline}} =
\mathbf{1014.89}\)).
Because CFI is calculated as: \[\text{CFI}_{\text{lavaan}} = 1 - \frac{62.23 - 24}{1014.89 - 36} = 1 - \frac{38.23}{978.89} = \mathbf{0.961}\] the inflated baseline \(\chi^2\) makes standard WLSMV CFI look artificially optimistic.
CatML Un-inflated Fit Indices in
umx: Savalei (2021) demonstrated that standard
WLSMV CFI suffers from artificial inflation when items have mixed
category counts. The catML approach implemented in
umx standardizes response metrics to unit diagonal (\(\text{diag}(\hat{\Sigma}^*) = 1.0\)),
eliminating category-level variance distortion. On mixed-category
models, catML reveals the true residual fit (\(\text{CFI} \approx 0.76 - 0.85\)),
preventing false confidence from baseline inflation.
umxSummary &
Writing Your PaperWhen you run umxSummary(model) on an ordinal WLS model,
umx outputs standard parameter tables and a statistical
note explaining the robust scaling parameters.
In the umxSummary() output note, you will see scaling
factors reported: \[\text{*Statistical Note*:
Ordinal WLS robust fit (Savalei (2021) categorical ML scaling (c_model =
1.815, c_null = 1.521).}\] \[\text{Unlike unscaled WLSMV which lacks this
calibration, Standard Hu & Bentler (1999) cutoffs apply (CFI >
.95, RMSEA < .06).}\] * Definition: \(c_{\text{model}}\) and \(c_{\text{null}}\) are the mean-and-variance
scaling correction multipliers applied to the model and baseline catML
discrepancies (\(XX_3\)). *
Interpretation: A value near \(1.0\) indicates that sample moment sampling
variability matches standard Maximum Likelihood expectations. Values
\(> 1.0\) adjust for overdispersion
and kurtosis inherent in categorical threshold and polychoric
correlation estimations.
uncertainty = "SE" vs
"RobustSE")umx are calculated
using the asymptotic covariance matrix \(\Gamma\) and WLS weight matrix \(W\) (the Generalized Method of Moments
sandwich estimator).uncertainty = "RobustSE" under WLS prints:
\[\text{*SEs*: WLS robust SEs (sandwich
estimator using asymptotic covariance of summary statistics).}\]
This confirms that parameter standard errors under WLS are robust to
non-normality by design.When reporting model fit in a journal article, report the robust catML fit indices directly alongside standard Hu & Bentler (1999) benchmark criteria:
“Model parameters were estimated using Diagonally Weighted Least Squares (DWLS) in
umx(Bates et al., 2026). Model fit was evaluated using robust Categorical Maximum Likelihood (catML) fit indices with mean-and-variance scaling adjustments (Savalei, 2021). The model demonstrated good fit to the data, \(\chi^2(24) = 42.52, p = .011\), \(\text{CFI} = .961\), \(\text{TLI} = .941\), and \(\text{RMSEA} = .047\) [\(90\%\) CI: \(.021, .073\)], satisfying standard benchmark criteria (\(\text{CFI} \ge .95\), \(\text{RMSEA} \le .06\); Hu & Bentler, 1999).”
| Metric | lavaan (WLSMV) |
umx (Savalei 2021 catML) |
Statistical Mechanism |
|---|---|---|---|
| Unscaled Discrepancy | \(42.99\) (\(\text{df}=24\)) | \(42.52\) (\(\text{df}=24\)) | Identical unscaled model fit |
| Scaling Factor | \(c = 0.736\) | \(c_3 = 1.815\) | Satorra–Bentler vs. catML correlation scaling |
| Reported \(\chi^2\) | \(62.23\) | \(42.52\) | lavaan inflates by \(1/0.736\); umx rescales \(77.16/1.815\) |
| Baseline Null \(\chi^2\) | \(1014.89\) (\(\text{df}=36\)) | \(120.4\) (\(\text{df}=36\)) | lavaan baseline is inflated by mixed
category weighting |
| CFI / TLI | \(0.961 / 0.941\) | \(0.758 / 0.637\) | lavaan CFI is baseline-inflated;
umx catML is un-inflated |
| Applicable Cutoffs | Non-standard | Hu & Bentler (1999) | Conventional cutoffs apply directly to umx
catML |
umx (Savalei catML)
vs. lavaan (WLSMV)umxSummary & Writing Your
Paper