M.Sc. Student, Computer Engineering
Istanbul Aydin University
I am a researcher working on the reliability and trustworthiness of large language models. My work is driven by a direct concern: LLMs are now accessible to everyone, yet their reasoning failures remain poorly understood and difficult to anticipate. I believe technology should genuinely help people, and that is only possible when the systems they rely on are transparent about what they can and cannot do.
My research focuses on understanding why LLMs fail on logical inference, predicting when they will fail, and designing interventions that make their behaviour more consistent and interpretable. Reliable, transparent AI is not a secondary concern, it is the foundation everything else has to be built on.
Research Interests: Large Language Models, Logical Reasoning, Modal Logic, Failure Prediction, Trustworthy AI, Natural Language Processing, Prompt Engineering
KI 2026: 49th German Conference on Artificial Intelligence, Bremen, Germany
Lecture Notes in Computer Science (LNAI, vol. 16830), Springer, Cham, pp. 317-323
F. Shahrokhshahi, F. Mohammadi
Investigated whether LLM failures on seven modal and conditional inference patterns are predictable from the structural properties of reasoning problems alone, without access to model outputs. An external classifier trained on 3,776 instances across nine models and three prompting strategies achieves AUC-ROC values of 0.69-0.93, with "must"-operator patterns proving most predictable. Results show that targeted prompting (LogiCue) and output-side reliability mechanisms serve complementary roles in LLM reasoning pipelines.
MeMo Workshop on Mechanistic Interpretability & Neuro-symbolic Approaches
University of Rome Tor Vergata · Accepted, June 2026 · Forthcoming in CEUR & arXiv
F. Shahrokhshahi
Explores explicit inference rule storage within MeMo's Correlation Matrix Memory as a mechanism for injecting structured logical reasoning capabilities into associative memory architectures.
ACLing 2025: 7th International Conference on AI in Computational Linguistics
Procedia Computer Science, vol. 275, pp. 484-492, Elsevier
F. Shahrokhshahi, F. Mohammadi, F. Sonmez
Developed a pattern-specific prompting methodology achieving 82.8% accuracy on challenging inference patterns, representing a 50.9 percentage point improvement over baseline approaches across nine state-of-the-art models including GPT-5, Claude Sonnet 4.5, and DeepSeek Reasoner.
Intelligent Academic Retrieval System (Nov 2024 - Feb 2025)
Balanced K-Means for Domain Discovery in Language Models (Feb 2025 - May 2025)
Earthquake Prediction Using Machine Learning (Under Preparation)
M.Sc. in Computer Engineering (Sep 2024 - present)
Istanbul Aydin University, Turkey
GPA: 4.0/4.0
B.Sc. in Robotic Engineering (Sep 2015 - Feb 2022)
Shahrood University of Technology, Iran