Publications
Empirical Inference
Deep Models and Optimization
Conference Paper
Scaling Behavior of Discrete Diffusion Language Models
von Rütte, D., Fluri, J., Pooladzandi, O., Schölkopf, B., Hofmann, T., Orvieto, A.
The Fourteenth International Conference on Learning Representations (ICLR), April 2026 (Published)
arXiv
URL
BibTeX
Empirical Inference
Robust Machine Learning
Conference Paper
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
Reizinger*, P., Mucsányi*, B., Guo*, S., Eysenbach, B., Schölkopf, B., Brendel, W.
The Fourteenth International Conference on Learning Representations (ICLR), April 2026, *equal contribution (Published)
arXiv
URL
BibTeX
Empirical Inference
Deep Models and Optimization
Conference Paper
Generalized Interpolating Discrete Diffusion
von Rütte, D., Fluri, J., Ding, Y., Orvieto, A., Schölkopf, B., Hofmann, T.
Proceedings of the 42nd International Conference on Machine Learning (ICML), 267:61810-61843, Proceedings of Machine Learning Research, (Editors: Singh, Aarti and Fazel, Maryam and Hsu, Daniel and Lacoste-Julien, Simon and Berkenkamp, Felix and Maharaj, Tegan and Wagstaff, Kiri and Zhu, Jerry), PMLR, International Conference on Machine Learning, July 2025 (Published)
arXiv
URL
BibTeX
Empirical Inference
Robust Machine Learning
Conference Paper
Cross-Entropy Is All You Need to Invert the Data Generating Process
Reizinger*, P., Bizeul*, A., Juhos*, A., Vogt, J. E., Balestriero, R., Brendel, W., Klindt, D.
The Thirteenth International Conference on Learning Representations (ICLR), April 2025, *Joint first authorship (Published)
arXiv
BibTeX
Empirical Inference
Robust Machine Learning
Conference Paper
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
Reizinger, P., Guo, S., Huszár, F., Schölkopf, B., Brendel, W.
The Thirteenth International Conference on Learning Representations (ICLR), April 2025 (Published)
arXiv
BibTeX
Empirical Inference
Robust Machine Learning
Conference Paper
Interaction Asymmetry: A General Principle for Learning Composable Abstractions
Brady, J., von Kügelgen, J., Lachapelle, S., Buchholz, S., Kipf*, T., Brendel*, W.
The Thirteenth International Conference on Learning Representations (ICLR), April 2025, *joint senior author (Published)
arXiv
BibTeX
Deep Models and Optimization
Conference Paper
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
Movahedi, S., Orvieto, A., Moosavi-Dezfooli, S.
In The Thirteenth International Conference on Learning Representations, ICLR 2025, The Thirteenth International Conference on Learning Representations, January 2025 (Accepted)
BibTeX
Robust Machine Learning
Conference Paper
Cross-Entropy Is All You Need To Invert the Data Generating Process
Reizinger, P., Bizeul, A., Juhos, A., Vogt, J. E., Balestriero, R., Brendel, W., Klindt, D.
In January 2025 (Published)
OpenReview
BibTeX
Robust Machine Learning
Conference Paper
Identifiable Exchangeable Mechanisms for Causal Structure and Representation Learning
Reizinger, P., Guo, S., Huszár, F., Schölkopf, B., Brendel, W.
In January 2025 (Published)
OpenReview
BibTeX
Robust Machine Learning
Conference Paper
In Search of Forgotten Domain Generalization
Mayilvahanan, P., Zimmermann, R. S., Wiedemer, T., Rusak, E., Juhos, A., Bethge, M., Brendel, W.
In January 2025 (Published)
OpenReview
BibTeX
Robust Machine Learning
Conference Paper
Interaction Asymmetry: A General Principle for Learning Composable Abstractions
Brady, J., von Kügelgen, J., Lachapelle, S., Buchholz, S., Kipf, T., Brendel, W.
In January 2025 (Published)
OpenReview
BibTeX
Deep Models and Optimization
Conference Paper
Using Shapley interactions to understand how models use structure
Divyansh Singhvi, D. M. A. E. R. J. I. P. N. S.
In Proceedings ACL, 1-20, Vienna Center, Association for Computational Linguistics (ACL 2025), 2025 (Accepted)
DOI
URL
BibTeX
Deep Models and Optimization
Article
An accelerated lyapunov function for Polyak’s Heavy-ball on convex quadratics
Orvieto, A.
Optimization Letters, 19:307-328, 2025 (Published)
DOI
URL
BibTeX
Deep Models and Optimization
Conference Paper
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
Monzio Compagnoni, E., Liu, T., Islamov, R., Proske, F. N., Orvieto, A., Lucchi, A.
In The Thirteenth International Conference on Learning Representations, ICLR 2025, The Thirteenth International Conference on Learning Representations, November 2024 (Accepted)
BibTeX
Robust Machine Learning
Article
Interaction Asymmetry: A General Principle for Learning Composable Abstractions
Brady, J., von Kügelgen, J., Lachapelle, S., Buchholz, S., Kipf, T., Brendel, W.
November 2024 (Submitted)
BibTeX
Deep Models and Optimization
Article
NIMBA: Towards Robust and Principled Processing of Point Clouds With SSMs
Köprücü, N., Okpekpe, D., Orvieto, A.
October 2024 (In preparation)
BibTeX
Deep Models and Optimization
Conference Paper
Loss Landscape Characterization of Neural Networks without Over-Parametrization
Islamov, R., Ajroldi, N., Orvieto, A., Lucchi, A.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, October 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
Zucchet, N., Orvieto, A.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, October 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Theoretical Foundations of Deep Selective State-Space Models
Muca Cirone, N., Orvieto, A., Walker, B., Salvi, C., Lyons, T.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, October 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
Sieber, J., Amo Alonso, C., Didier, A., Zeilinger, M., Orvieto, A.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, October 2024 (Published)
URL
BibTeX
Robust Machine Learning
Conference Paper
Measuring Per-Unit Interpretability at Scale Without Humans
Klindt, D., Zimmermann, R., Brendel, W.
In September 2024 (Published)
OpenReview
BibTeX
Robust Machine Learning
Conference Paper
Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts
Mészáros, A., Ujváry, S., Brendel, W., Reizinger, P., Huszár, F.
In September 2024 (Published)
ArXiv
BibTeX
Deep Models and Optimization
Article
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
Orvieto, A., Xiao, L.
July 2024 (Submitted)
BibTeX
Robust Machine Learning
Conference Paper
InfoNCE: Identifying the Gap Between Theory and Practice
Rusak, E., Reizinger, P., Juhos, A., Bringmann, O., Zimmermann, R. S., Brendel, W.
In July 2024 (Published)
BibTeX
Empirical Inference
Robust Machine Learning
Conference Paper
Position: Understanding LLMs Requires More Than Statistical Generalization
Reizinger, P., Ujváry, S., Mészáros, A., Kerekes, A., Brendel, W., Huszár, F.
Proceedings of the 41st International Conference on Machine Learning (ICML), 235:42365-42390, Proceedings of Machine Learning Research, (Editors: Salakhutdinov, Ruslan and Kolter, Zico and Heller, Katherine and Weller, Adrian and Oliver, Nuria and Scarlett, Jonathan and Berkenkamp, Felix), PMLR, July 2024 (Published)
arXiv
URL
BibTeX
Deep Models and Optimization
Article
An accelerated Lyapunov function for Polyak’s Heavy-ball on convex quadratics
Orvieto, A.
Optimization Letters, June 2024 (Published)
BibTeX
Deep Models and Optimization
Article
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
Yi Meng, S., Orvieto, A., Yiming Cao, D., De Sa, C.
June 2024 (Submitted)
BibTeX
Deep Models and Optimization
Conference Paper
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
Orvieto, A., De, S., Gulcehre, C., Pascanu, R., Smith, S. L.
In Proceedings of Machine Learning Research, Proceedings of the Forty-First International Conference on Machine Learning , Forty-First International Conference on Machine Learning , June 2024 (Published)
URL
BibTeX
Robust Machine Learning
Conference Paper
Does CLIP’s Generalization Performance Mainly Stem from High Train-Test Similarity?
Mayilvahanan, P., Wiedemer, T., Rusak, E., Bethge, M., Brendel, W.
In June 2024 (Published)
ArXiv
BibTeX
Robust Machine Learning
Conference Paper
Don’t trust your eyes: on the (un) reliability of feature visualizations
Geirhos, R., Zimmermann, R. S., Bilodeau, B., Brendel, W., Kim, B.
In June 2024 (Published)
ArXiv
BibTeX
Robust Machine Learning
Article
Translational symmetry in convolutions with localized kernels causes an implicit bias toward high frequency adversarial examples
Caro, J. O., Ju, Y., Pyle, R., Dey, S., Brendel, W., Anselmi, F., Patel, A. B.
Frontiers in Computational Neuroscience, 18:1387077, June 2024 (Published)
Frontiers in Computational Neuroscience
BibTeX
Robust Machine Learning
Conference Paper
An interventional perspective on identifiability in gaussian lti systems with independent component analysis
Rajendran, G., Reizinger, P., Brendel, W., Ravikumar, P. K.
41-70, Causal Learning and Reasoning, March 2024 (Published)
PMLR
BibTeX
Deep Models and Optimization
Conference Paper
SDEs for Minimax Optimization
Monzio Compagnoni, E., Orvieto, A., Kersting, H., Proske, F., Lucchi, A.
PMLR, AISTATS, February 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Recurrent Distance Filtering for Graph Representation Learning
Ding, Y., Orvieto, A., He, B., Hofmann, T.
In PMLR, ICML, January 2024 (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Super Consistency of Neural Network Landscapes and Learning Rate Transfer
Noci, L., Meterez, A., Hofmann, T., Orvieto, A.
In Proceedings of the Thirty-Eighth Annual Conference on Neural Information Processing Systems, Thirty-Eighth Annual Conference on Neural Information Processing Systems, January 2024 (Published)
URL
BibTeX
Robust Machine Learning
Conference Paper
Effective pruning of web-scale datasets based on complexity of concept clusters
Abbas, A., Rusak, E., Tirumala, K., Brendel, W., Chaudhuri, K., Morcos, A. S.
In January 2024 (Published)
ArXiv
BibTeX
Robust Machine Learning
Conference Paper
Provable Compositional Generalization for Object-Centric Learning
Wiedemer, T., Brady, J., Panfilov, A., Juhos, A., Bethge, M., Brendel, W.
In October 2023 (Published)
ArXiv
BibTeX
Robust Machine Learning
Conference Paper
Scale Alone Does not Improve Mechanistic Interpretability in Vision Models
Zimmermann, R. S., Klein, T., Brendel, W.
In Advances in Neural Information Processing Systems 36 (NeurIPS 2023), 57876 - 57907, Curran Associates Inc., NeurIPS, October 2023 (Published)
NeurIPS Proceedings
DOI
URL
BibTeX
Robust Machine Learning
Conference Paper
Compositional Generalization from First Principles
Wiedemer, T., Mayilvahanan, P., Bethge, M., Brendel, W.
In July 2023 (Published)
NeurIPS Proceedings
BibTeX
Empirical Inference
Robust Machine Learning
Conference Paper
Desiderata for Representation Learning from Identifiability, Disentanglement, and Group-Structuredness
Keurti, H., Reizinger, P., Schölkopf, B., Brendel, W.
2nd Annual Topology, Algebra, and Geometry in Machine Learning (TAG) at ICML 2023, July 2023 (Published)
URL
BibTeX
Empirical Inference
Robust Machine Learning
Conference Paper
Provably Learning Object-Centric Representations
Brady*, J., Zimmermann*, R. S., Sharma, Y., Schölkopf, B., von Kügelen, J., Brendel, W.
Proceedings of the 40th International Conference on Machine Learning (ICML), 202:3038-3062, Proceedings of Machine Learning Research, (Editors: A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato and J. Scarlett), JMLR, Cambridge, MA, July 2023, *equal contribution (Published)
URL
BibTeX
Deep Models and Optimization
Conference Paper
Resurrecting Recurrent Neural Networks for Long Sequences
Orvieto, A., Smith, S. L., Gu, A., Fernando, A., Gulcehre, C., Pascanu, R., De, S.
In Proceedings of the Eleventh International Conference on Learning Representations, ICLR, June 2023 (Published)
URL
BibTeX
Empirical Inference
Robust Machine Learning
Article
Jacobian-based Causal Discovery with Nonlinear ICA
Reizinger, P., Sharma, Y., Bethge, M., Schölkopf, B., Huszár, F., Brendel, W.
Transactions on Machine Learning Research, April 2023 (Published)
URL
BibTeX