PapersWithELO
← ICLR 2024 leaderboard

Learning Latent Structural Causal Models

Jithendaraa Subramanian, Yashas Annadani, Tristan Deleu, Ivaxi Sheth, Nan Rosemary Ke, Stefan Bauer, Derek Nowrouzezahrai, Samira Ebrahimi Kahou

probabilistic methodsBayesian Causal DiscoveryLatent variable models
53.40100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
68.50100
Mimo
band ≈ ±20 pct pts (from σ = 0.41)
44.20100
DeepSeek
band ≈ ±19 pct pts (from σ = 0.39)

OpenReview ground truth

Rejected

TL;DR — Bayesian inference over latent structural causal models from low-level data under random, known interventions for linear Gaussian additive noise SCMs. Such a model also performs image generation from unseen interventions.

Abstract

Causal learning has long concerned itself with the recovery of underlying causal mechanisms. Such causal modelling enables better explanations of out-of-distribution data. Prior works on causal learning assume that the causal variables are given. However, in machine learning tasks, one often operates on low-level data like image pixels or high-dimensional vectors. In such settings, the entire Structural Causal Model (SCM) -- structure, parameters, \textit{and} high-level causal variables -- is latent and needs to be learnt from low-level data. We treat this problem as Bayesian inference of the latent SCM, given low-level data. We present BIOLS, a tractable approximate inference method which performs joint inference over the causal variables, structure and parameters of the latent SCM from known interventions. Experiments are performed on synthetic datasets and a causal benchmark image dataset to demonstrate the efficacy of our approach. We also demonstrate the ability of BIOLS to generate images from unseen interventional distributions.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 38)