RIML Lab
СтатистикаRobust and Interpretable Machine Learning Lab, Prof. Mohammad Hossein Rohban, Sharif University of Technology https://youtube.com/@rimllab twitter.com/MhRohban https://www.linkedin.com/company/robust-and-interpretable-machine-learning-lab/
- Последний пост
- 10 авг.
- Последнее чтение
- 01:39
- Постов за неделю
- 1
- Всего постов
- 20
- Тип
- открытый
- Язык
- английский
- Категория
- Образование
- В каталоге с
- 12 авг.
- 1/24сутки в ленте
- 910
- 1/48двое суток
- 1 042
- 1/72трое суток
- 1 124
Оценка по просмотрам недавних постов: пост набирает почти всё за первые сутки.
Посты
🔐 ML Security Journal Club ✅ This Week's Presentation: 🔹 Title: How Jailbreaks Evade, but Do Not Erase, LLM Safety Mechanisms 🔸 Presenter: Javad Hezareh 🌀 Abstract: This paper investigates the internal mechanisms of Large Language Models (LLMs) during successful jailbreak attacks. The authors provide mechanistic evidence that jailbreaks do not comprehensively eliminate an LLM's safety features; instead, they selectively suppress specific components to bypass refusal mechanisms, leaving other robust internal safety representations intact. To validate the utility of these mechanistic insights, the authors developed a training-free harmful-content detector. By reading the robust internal activations without any model training, this detector achieves competitive aggregate performance and strong adversarial robustness on safety-eval benchmarks. 📄 Paper: Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Session Details: * 📅 Date: Tuesday, Aug 11 * 🕒 Time: 14:00 - 15:00 * 🌐 Location: Online at vc.sharif.edu/ch/rohban We look forward to your participation! ✌️
📈 Generative Modeling — Scaling, Multimodality & End-to-End Generation ✅ This Week's Presentation: 🔹 Title: Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation 🔸 Presenter: Amir Qeysarbeigi 🌀 Abstract: Modern generative models typically handle multimodal distributions by factorizing the generation process into multiple steps, as in autoregressive and diffusion models. While this enables high-quality generation, it creates a mismatch between training and inference and prevents fully end-to-end generation. In this work, we introduce Explorative Modeling (XM), a new paradigm that instead factorizes the training process by exploring multiple candidate generations and training on the best-matching one. This exploration increases generative expressivity, allowing models to capture more modes of multimodal distributions without relying solely on generation factorization. The paper demonstrates that exploration acts as a third pretraining scaling axis, alongside model parameters and data, improving efficiency across image, video, and language generation. Increasing exploration improves FLOP, sample, and parameter efficiency, with gains that become larger as models and datasets scale. Furthermore, by moving the burden of multimodality from inference-time generation steps to training-time exploration, Explorative Modeling enables end-to-end generative models that can achieve performance comparable to diffusion-based approaches with dramatically fewer inference steps. 📄 Article: *Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation* Session Details: We will first review the limitations of conventional reconstructive generative models and introduce the concept of generative expressivity as a fundamental bottleneck in multimodal generation. Then, we will explore how Explorative Modeling replaces generation factorization with training-time exploration, including the Forward and Reverse XM formulations. Finally, we will examine how exploration serves as a new scaling axis, improves efficiency across multiple modalities, and enables end-to-end generation with substantially fewer inference steps. - 📅 Date: Monday (دوشنبه) - 🕒 Time: 17:00 - 18:00 - 🌐 **Location: https://vc.sharif.edu/rohban** (Online only) We look forward to your participation! ✌️
📢 Research Collaboration in Quantitative Finance at RIML We are seeking motivated students interested in quantitative finance, stochastic modeling, machine learning, and portfolio optimization. Selected researchers will work under the supervision of Dr. Rohban and collaborate with international professors and researchers affiliated with the University of Manchester, the Alan Turing Institute, Virginia Tech, and the Technical University of Munich 🔬 The following four research directions are available: 1️⃣ Reinforcement Learning in High-Frequency Market Making Study the trade-off between time discretization, learning accuracy, and sample complexity in single- and multi-agent market making, including convergence to continuous-time optimal policies and Nash equilibria. 🔗 Paper: https://arxiv.org/abs/2407.21025 2️⃣ Stochastic Optimal Control for Multi-Asset Market Making Develop scalable closed-form approximations to the Hamilton–Jacobi equations of multi-asset market-making models, enabling interpretable near-optimal quotes under correlated prices and portfolio-wide inventory risk. 🔗 Paper: https://arxiv.org/abs/1810.04383 3️⃣ Diffolio: A Diffusion Model for Multivariate Probabilistic Financial Time-Series Forecasting and Portfolio Construction Model the conditional joint distribution of future asset returns using hierarchical asset-level and market-level attention, and use the generated scenarios for risk-aware portfolio construction. 🔗 Paper: https://arxiv.org/abs/2511.07014 4️⃣ Structured Filtering for Jump-Diffusion Time Series Forecasting to infer hidden market states from partially observed jump-diffusion data and produce calibrated probabilistic forecasts of continuous movements and abrupt price shocks. 🔗 Paper: https://arxiv.org/abs/2605.24548 ✉️ Interested candidates are invited to send their CV to: alirezanourimath@gmail.com
آزمایشگاه RIML تحت نظارت دکتر رهبان در حال راهاندازی یک ژورنالکلاب پیرامون یادگیری تقویتی چندعاملی (Multi-Agent RL) بر پایهی کتاب Albrecht با چشمانداز حرکت به سمت کار پژوهشی جدی در این حوزه است. در صورت علاقهمندی به این مسیر، خواهشمندست این فرم را پر کنید. *: جلسههای ژورنالکلاب از این هفته آغاز میشود. در صورتی که پرسش یا ابهامی در این زمینه دارید با شناسهی زیر در تلگرام ارتباط بگیرید: @Moein_Salimi
🔐 LLM Faithfulness Journal Club ✅ This Week's Presentation: 🔹 Title: Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations 🔸 Presenter: [Farzan Rahmani](https://www.linkedin.com/in/farzan-rahmani-51128b201) 🌀 Abstract: When large language models explain their decisions, their explanations may sound convincing—but do they faithfully reflect the model's true reasoning? This paper presents a comprehensive counterfactual analysis of self-explanation faithfulness across 75 models from 13 model families. It investigates the tradeoff between concise and comprehensive explanations, introduces two new evaluation metrics (phi-CCT and F-AUROC), and studies how explanation verbosity influences faithfulness measurements. The results reveal a clear scaling trend: larger and more capable language models consistently produce more faithful self-explanations, providing valuable insights into the relationship between model scale, explanation quality, and AI safety. 📄 Paper: [Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations](https://arxiv.org/abs/2503.13445) Session Details: * 📅 Date: Wednesday (چهارشنبه) * 🕑 Time: 2:00 - 3:00 PM * 🌐 Location: Online at http://vc.sharif.edu/ch/rohban We look forward to your participation! ✌️
🤖 RL Journal Club ✅ This Week's Presentation: 🔹 Title: Is Reinforcement Learning Really Harder Than Bandits? 🔸 Presenter: Arshia Gharooni 🌀 Abstract: Episodic reinforcement learning, despite having longer planning horizons, presents little additional sample complexity difficulty compared to contextual bandits, with the proposed Monotonic Value Propagation (MVP) algorithm achieving near-optimal regret bounds. The MVP algorithm utilizes a simplified, variance-aware bonus to achieve superior performance, offering an exponential improvement in horizon dependency and sample efficiency over previous state-of-the-art methods. 📄 Paper: Is Reinforcement Learning More Difficult Than Bandits? A Near-optimal Algorithm Escaping the Curse of Horizon Session Details: - 📅 Date: Wednesday (چهارشنبه) - 🕒 Time: 17:00 - 18:00 - 🌐 Location: Online at http://vc.sharif.edu/ch/rohban We look forward to your participation! ✌️
🤖 RL Journal Club ✅ This Week's Presentation: 🔹 Title: Test Time Exploration to Achieve Generalization in Zero-Shot RL 🔸 Presenter: Alireza Farajtabrizi 🌀 Abstract: This paper studies zero-shot generalization in reinforcement learning, where an agent is trained on a set of tasks but must perform well on unseen test environments. The authors argue that standard reward-maximizing RL agents can overfit to training tasks, especially in environments where simple invariance-based methods fail. Their key insight is that exploration behavior is harder to memorize than reward-seeking behavior and can therefore generalize better. To build on this idea, the paper introduces Explore to Generalize (ExpGen), an algorithm that combines a maximum-entropy exploration policy with an ensemble of reward-seeking agents. At test time, when the ensemble agrees on an action, the agent exploits that decision; when the ensemble is uncertain, the agent switches to the exploration policy to reach new parts of the state space. Experiments on ProcGen show strong improvements on challenging tasks such as Maze and Heist, setting new state-of-the-art results in several zero-shot RL settings. 📄 Paper: Explore to Generalize in Zero-Shot RL (NeurIPS 2023) Session Details: * 📅 Date: Tuesday سهشنبه * 🕒 Time: 15:30 - 16:30 * 🌐 Location: Online at vc.sharif.edu/ch/rohban (http://vc.sharif.edu/ch/rohban) We look forward to your participation! ✌️
🚀 Open Research Position: Visual Reasoning in Large Vision-Language Models (LVLMs) We are looking for motivated students to join our research on visual reasoning in Large Vision-Language Models (LVLMs) at RIML Lab. 🔍 Project Description Large Vision-Language Models have achieved remarkable performance across a wide range of multimodal tasks. However, their ability to perform complex visual reasoning remains an open challenge. This research focuses on understanding, evaluating, and improving the reasoning capabilities of LVLMs, including multi-step reasoning, visual grounding, and reasoning over complex visual scenes. 📄 Relevant Papers Question Aware Vision Transformer for Multimodal Reasoning Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning 🔹 Must-Have Requirements Strong Python programming skills Knowledge of deep learning and machine learning fundamentals Hands-on experience with PyTorch Familiarity with Vision-Language Models or Large Language Models Strong research interest and willingness to learn Ready to start immediately ⏳ Workload Commitment: At least 20 hours per week 📌 Note: Filling out this form does not guarantee acceptance. Only shortlisted candidates will be contacted via email. 🔗 Apply here: Form 💬 Telegram: @Arianaghamohseni @RIMLLab #research_position #ML_research #VisionLanguageModels #MultimodalAI #VisualReasoning #DeepLearning
🔐 LLM Faithfulness Journal Club ✅ This Week's Presentation: 🔹 Title: Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety 🔸 Presenter: [Farzan Rahmani](https://www.linkedin.com/in/farzan-rahmani-51128b201) 🌀 Abstract: AI systems that “think” in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability. 📄 Paper: [Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety](https://arxiv.org/abs/2507.11473v2) Session Details: - 📅 Date: Wednesday (چهارشنبه) - 🕒 Time: 10:00 - 11:00 AM - 🌐 Location: Online at http://vc.sharif.edu/ch/rohban We look forward to your participation! ✌️
🔐 ML Security Journal Club ✅ This Week's Presentation: 🔹 Title: A Machine Unlearning Approach to Safety Alignment 🔸 Presenter: Arian Komaei 🌀 Abstract: The paper identifies a fundamental limitation in current vision language model (VLM) alignment called the "safety mirage." Traditional supervised safety fine-tuning often reinforces superficial textual patterns rather than deep harm mitigation, leaving models vulnerable to simple one-word attacks and causing "over-prudence" (unnecessary rejections of benign queries). To address this, the authors propose Machine Unlearning (MU) as a superior alternative. Unlike standard fine-tuning, MU directly removes harmful knowledge and avoids biased feature-label mappings. Extensive evaluations show that MU-based alignment reduces attack success rates by up to 60.27% and cuts unnecessary rejections by over 84.20%, all while preserving the model's general capabilities. 📄 Paper: Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning Session Details: * 📅 Date: Thursday پنج شنبه * 🕒 Time: 9:00 - 10:00 AM * 🌐 Location: Online at vc.sharif.edu/ch/rohban We look forward to your participation! ✌️
🔘 Open Research Positions: Shortcut Learning and Spurious Correlation We are looking for motivated students to join our research projects. 🔍 Project Description Shortcut learning occurs when models rely on spurious or overly simple patterns instead of learning the intended underlying task. Our research aims to understand why and how shortcuts emerge, and how they can be identified, analyzed, and mitigated. We offer three research assistant positions aligned with the following directions, co-advised by Dr. Rohban and Dr. Soleymani: 🔹 1. Understanding the Implicit Effects of Inductive Biases on Shortcut Learning This project investigates how training-related inductive biases, such as batch size, learning rate, and loss function, influence the emergence of shortcut learning and group robustness. [1] The Silent Helper: How Implicit Regularization Enhances Group Robustness (HiLD Workshop ICLR 2025) [2] On the Role of Implicit Regularization of Stochastic Gradient Descent in Group Robustness (ICLR 2026) 🔗 Apply here: Google Form 🔹 2. Understanding Shortcut Learning Through the Lens of Task Arithmetic It has been observed that shortcut features are often learned early during training. By analyzing the trajectory of weight evolution, we aim to identify shortcut task directions in parameter space and leverage this understanding to mitigate shortcut reliance. [3] Editing Models with Task Arithmetic (ICLR 2023) 🔗 Apply here: Google Form 🔹 3. Investigating Shortcut Learning in Large-Scale Models Shortcut-based generalization failures were first studied in linear models and simple settings. In this project, we aim to understand how these biases propagate to large-scale models, including language models, and explore strategies to mitigate them. [4] Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models (EMNLP 2024) [5] Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful (NeurIPS 2025) 🔗 Apply here: Google Form 📌 Important Notes: Please carefully read the Internship Ethics and Guidelines before submitting your application. Submitting the application form does not guarantee acceptance. Only shortlisted candidates will be contacted via email by March 15.
دانشجویانی که علاقه مند به دستیاری آموزشی درس یادگیری تقویتی دکتر رهبان در نیم سال جاری هستند لطفا فرم زیر را پر کنند. https://docs.google.com/forms/d/e/1FAIpQLSe2rPGlxxTQDzDnt6PB_hRAtBa64_OLmWLE_dUnPqQlAfpUqQ/viewform?usp=dialog
We invite interested collaborators to join an ongoing research project that aims to redefine attention mechanisms in large language models, drawing inspiration from the concept of consciousness. The goal of this work is to produce results suitable for submission to NeurIPS 2026. This research is conducted under the supervision of Dr. Rohban and is led by his PhD student, Hamidreza Akbari. Motivated undergraduate students with a strong background or interest in deep learning/LLMs and related fields are encouraged to reach out to @hamidrakbari to explore potential collaboration opportunities.
دانشجویانی که علاقهمند به دستیاری آموزشی درس سیستم ۲ (دکتر رهبان، دکتر سلیمانی و آقای سمیعی) هستند، خواهشمند است فرم زیر را تکمیل نمایند: تکمیل فرم
📢 Open RA Positions: Generative Models (RIML & TSAIL Labs) Supervisors: Dr. Rohban & Dr. Sadeghzadeh (Sharif University of Technology) Focus: Robustness, Interpretability, and Trustworthiness in Generative Models. Goal: NeurIPS/ICLR submissions (4-month timeline). ✅ Requirements: - Strong background in ML/AI and Generative Models - Proficiency in Python/PyTorch - Commitment: 20–30 hours/week ⚠️Note: We do not accept applicants who currently have a full-time job or those who are students with a part-time job or have other research positions elsewhere. 📚 Background Reading - http://arxiv.org/abs/2410.15618 - http://arxiv.org/abs/2305.10120 👉 Apply Here
Hey folks! 👋 Our team is working on foundation models, with a strong focus on interpretability and reliability. We’re opening applications for new team members who are excited to learn, take ownership of tasks, and contribute consistently to solid, hands-on research and experimentation. Highly motivated bachelor’s students and junior undergraduates are especially encouraged to apply. Selection prioritizes motivation, dedication, perseverance, strong fundamentals, and consistent follow-through over grades. Experience with PyTorch is a major plus. Ideal candidates are comfortable implementing clean, reproducible experiments and communicating progress clearly. 👉 Apply Here. Looking forward to doing deep, high-impact, and fun science together! 🚀🥰🔥
🪢 Compositional Learning Journal Club Join us this week for a fascinating dive into how multimodal language models can think visually by drawing — mimicking a human’s use of sketches to guide reasoning and solve complex tasks. 🌟 This Week's Presentation 📄 Paper: Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models 🧠 Abstract: Multimodal LLMs are strong at visual reasoning, but they typically rely on text-only intermediate steps. This paper introduces Visual Sketchpad, which gives MLLMs a lightweight drawing interface (e.g., lines, boxes, marks) so they can create visual intermediate steps while reasoning—similar to how humans sketch when solving problems. By integrating these sketch actions (and optionally leveraging vision modules during sketching), the approach improves performance across a wide range of tasks, including math/geometry, graphs, and spatial reasoning. 🎙 Presenter: Amir Kasaei Session Details: - 📅 Date: Tuesday, December 30 - 🕒 Time: 3:00 - 4:00 PM - 🌐 Location: Online at vc.sharif.edu/ch/rohban We look forward to your participation! ✌️
🪢 Compositional Learning Journal Club Join us this week for a critical exploration of robustness in Visual Question Answering systems and the broader implications for visual–language model reliability. We’ll analyze how even subtle, meaning-preserving changes to inputs can destabilize model outputs and discuss what this means for future evaluation and model design. 🌟 This Week's Presentation 📄 Paper: Questioning the Stability of Visual Question Answering 🧠 Abstract: Modern Visual Language Models (VLMs) have achieved impressive performance on a wide range of visual reasoning tasks, yet fundamental questions remain about their robustness to benign input perturbations. This paper presents the first large-scale, systematic study of how VLMs respond to small, meaning-preserving changes—such as pixel shifts, light geometric transformations, padded rescaling, paraphrasing, and multilingual rewrites—that do not change the true semantics of an image–question pair. Across multiple datasets and models, the authors find that minor visual or textual perturbations frequently lead to different predicted answers, even for state-of-the-art systems like GPT-4o and Gemini 2.0 Flash. They also show that stability under perturbations correlates strongly with correctness, and that the stability patterns of small open-source models can be used to predict when larger models will fail. In this session, we’ll discuss: • What kinds of input changes most disrupt VQA predictions. • How stability can serve as a proxy for reliability and model confidence. • Implications for evaluation benchmarks and future model development. 🎙 Presenter: Amir Kasaei Session Details: - 📅 Date: Tuesday, December 23rd - 🕒 Time: 3:00 PM - 4:00 PM - 🌐 Location: Online at vc.sharif.edu/ch/rohban We look forward to your participation! ✌️
🔐 ML Security Journal Club ✅ This Week's Presentation: 🔹 Title: Unlearning diffusion models 🔸 Presenter: Arian Komaei 🌀 Abstract: The paper argues that modern multi-stage training pipelines create a fundamental obstacle for machine unlearning. The ideal goal—Retrain Equivalence—is for an unlearned model to behave exactly like one retrained from scratch without the forgotten data. But the authors show, both theoretically and empirically, that this is often impossible: once training happens in multiple stages with different data and objectives, the model’s behavior becomes path-dependent. That means the order of training steps permanently affects how unlearning works, and “local” unlearning methods that only use gradients from the forget set can’t universally reach Retrain Equivalence. Experiments on Llama and Qwen models (1B–14B) confirm strong divergence: the same data but different training orders lead to very different unlearning outcomes, with accuracy dropping by 20% across paths. Some training paths also produce models that are inherently harder to unlearn. Because multi-stage training is now standard and training histories are often unavailable, the paper concludes that Retrain Equivalence is the wrong target and the field needs to rethink what machine unlearning should actually aim for. 📄 Paper: ON THE IMPOSSIBILITY OF RETRAIN EQUIVALENCE IN MACHINE UNLEARNING Session Details: - 📅 Date: Sunday - 🕒 Time: 3:30 - 4:30 PM - 🌐 Location: Online at vc.sharif.edu/ch/rohban We look forward to your participation! ✌️
🪢 Compositional Learning Journal Club Join us this week for a deep dive into how CLIP actually represents multiple objects in an image—and where it silently goes wrong. We’ll look at subtle biases in both text and image encoders, how they interact with caption structure and object size, and what this means for downstream multimodal models and text-to-image generation. 🌟 This Week's Presentation 📌 Title: CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation 🧠 Abstract: Contrastive Language–Image Pre-training (CLIP) has become a workhorse for zero-shot classification and many vision–language tasks, but its behavior in complex scenes with multiple objects is far from fully understood. This session focuses on a systematic study of CLIP in controlled multi-object setups using ComCO, a dedicated dataset designed to probe how CLIP’s encoders handle object combinations and compositional structure. We will discuss evidence that: - The text encoder tends to over-focus on the first-mentioned object in a caption. - The image encoder tends to favor larger objects in the scene. - Small changes such as swapping token order or resizing objects can cause sharp drops in image–text matching and retrieval performance across multiple CLIP variants. 📄 Paper: CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation 🎙 Presenter: Dr MH Rohban Session Details: - 📅 Date: Tuesday, December 2nd - 🕒 Time: 3:00 PM - 4:00 PM - 🌐 Location: Online at vc.sharif.edu/rohban We look forward to your participation! ✌️