Offene Abschlussarbeiten


Uncovering Geometric Primitives in Object Representations

Vision Models achieve remarkable accuracy in categorizing objects, yet it remains unclear if these successes are driven by superficial texture matching or a deeper structural understanding of the world. Determining whether a model’s internal representation of a „table“ is grounded in the abstract concept of a „rectangle“ is essential for developing AI that perceives the world through human-like geometric reasoning. This research introduces the concept of approximate probes to quantify the geometric transfer between pure mathematical ideals and complex real-world entities. These probes are then deployed across the layers of a pre-trained vision transformer to evaluate if the model successfully recognizes the underlying geometry of real-world photographs without ever being trained on them.

References

Tenney, I., Das, D., & Pavlick, E. (2019). BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th annual meeting of the association for computational linguistics (pp. 4593-4601).

Arnold, S., & Gröbner, R. (2026). Locating and Editing Figure-Ground Organization in Vision Transformers. arXiv preprint arXiv:2603.06407.


Mechanistic Interpretability of Syntactic Reduction

Natural language frequently employs abbreviated structures, such as verb contractions (e.g., isn’t) and the omission of subordinating conjunctions. This thesis aims to dissect the shared functional mechanisms behind morphological contractions and that-deletion within language models using logit lens and activation patching.

References

https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens

Meng, K., Bau, D., Andonian, A., & Belinkov, Y. (2022). Locating and editing factual associations in gpt. Advances in neural information processing systems35, 17359-17372.

Arnold, S., & Gröbner, R. (2025,). Steering Prepositional Phrases in Language Models: A Case of with-headed Adjectival and Adverbial Complements in Gemma-2. In Proceedings of the 8th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP (pp. 69-78).


Uncovering Prompt Engineering Competency: Learning Effects in Wild Human–Agent Dialogs

Users frequently adapt their linguistic style over time when interacting with conversational AI, yet the trajectory remains largely unquantified at scale. This thesis investigates longitudinal skill acquisition by analyzing large-scale user-agent interaction logs from the WildChat corpus. By leveraging expected utility norms for English expressions, this research will track how prompt efficiency evolve across active user sessions, examining whether sustained interaction leads to systematic prompt refinement or persistent behavioral patterns.

References

Koulouri, T., Lauria, S., & Macredie, R. D. (2016). Do (and say) as I say: Linguistic adaptation in human–computer dialogs. Human–Computer Interaction31(1), 59-95.

Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., & Deng, Y. (2024). Wildchat: 1m chatgpt interaction logs in the wild. arXiv preprint arXiv:2405.01470.

Wang, A., Brysbaert, M., & Günther, F. (2026). Adding volition to word processing: Expected utility norms for 80,000 English words and multiword expressions. Behavior Research Methods58(7), 176.


Decoding Visual Power Structures in Text-to-Image Representations

Text-to-image models achieve remarkable fidelity in synthesizing complex scenes, yet it remains unclear how these architectures capture and project abstract social constructs, such as hierarchical power dynamics, into the visual domain. Determining whether a model’s generated depiction of an asymmetric relationship—such as a „CEO and an assistant“—is grounded in established socio-visual parameters is essential for understanding how artificial intelligence manifests human societal structures. This research systematically derives a catalog of historical and cultural power representations from literature, analyzing structural variables like spatial centrality or the geometric vectors and angles of human posture. Using this foundational framework, targeted prompts will be designed to qualitatively evaluate how generative models visually interpret and map the abstract variable of power imbalances without explicit spatial instructions.

References

Yixin Wan and Kai-Wei Chang. (2025). The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, pages 9174–9190, Vienna, Austria. Association for Computational Linguistics.