Lecturer Teaching Fellows Program · 2025–26
Reclaiming Authentic Self-Reflection
Does a restructured journaling assignment change how MBA students write? Research comparing Fall 2025 and Spring 2026 cohorts at UC Berkeley Haas.
Jennifer Caleshu · Lecturer Teaching Fellow 2025–26 · Download full poster (PDF)
Context & Motivation
How might we create self-reflection assignments that encourage deep reflection rather than shallow or inauthentic (i.e. AI generated) essays? Many of the leadership courses I teach at the Haas School of Business school primarily use self-reflection essays as the summative assessment. Over the past 10 years, I've seen these essays become progressively less reflective and generic. I suspect a variety of reasons, including the decline of liberal arts degrees, the growth of a more transactional workplace environment (hybrid/remote workplaces offer less informal human connection), and the recent impact of generative AI.
The original assignment asked students to "keep an ongoing journal during and after class" and attach at least three pages as an appendix to the final paper. The assignment included 20+ reflection questions, but no structural requirements. The final paper was submitted via the Rumi tool in bCourses, to capture behavioral data about how students were writing.
For Spring 2026, the assignment was redesigned to require three distinct, structured self-reflection entries (submitted separately from the final paper), with the option to complete them as written journals or recorded peer-dialogue transcripts.
Research Questions
Did the redesigned assignment change how students wrote the final paper?
Did it change the substance and depth of journal reflection?
Did the transcript option yield comparable engagement to written journals?
Research Methods
Cohort comparison
Fall 2025 (n=217) vs. Spring 2026 (n=109) final-paper Rumi process metrics compared across journal-type choices (written, transcript, combination). Section and timing checks used as bias controls.
Qualitative coding
AI (Claude) qualitative coding review of a stratified sample of 30 journal entries (15 per semester), coded for specificity, integration with course concepts, and reflection depth.
Results: Process Metrics
Rumi process metrics for final papers, comparing Fall 2025 (n=217) and Spring 2026 (n=109). Significance levels: *p<.05, **p<.01, ***p<.001. Final paper grades held constant by design—graded to a 3.65 GPA target.

| Metric | Fall 2025 | Spring 2026 | Change |
|---|---|---|---|
| Writing Duration(sec) | 4,852 | 4,240 | -13% |
| Active Time(sec) | 8,233 | 6,323 | -23%** |
| Total Deletes | 883 | 611 | -31% |
| Total Time(sec) | 11,012 | 8,068 | -27%** |
| Text Revised %(%) | 31.5 | 25.9 | -18%* |
| Writing Speed(WPM) | 19.7 | 21.9 | +11% |
| AI Tool Usage(calls) | 5.1 | 1.7 | -67%*** |
Results: Journal Quality
Specificity and course-concept integration rose modestly
Qualitative coding of the 30-journal sample showed that Spring entries were more likely to name specific class exercises, reference particular readings, and connect personal experiences to named leadership frameworks. However, the sample size (15 per semester) is too small to generalize, and a single coder limits reliability.
Whether reflection itself deepened could not be determined
Depth of reflection—genuine insight vs. surface-level description—is harder to code reliably than specificity. With 15 journals and a single AI coder, no confident conclusions can be drawn about whether the redesign changed how deeply students reflected, as opposed to how they organized and expressed that reflection.
Format choice (journal vs. transcript) did not predict performance
Students who submitted dialogue transcripts showed comparable Rumi process metrics and qualitative coding results to those who wrote journals. Format flexibility widened access without compromising process or content quality.
Conclusions
▸ Scaffolding works at the level of process and structure
Students who completed structured pre-reflection arrived at the final paper with their thinking already organized. They wrote more efficiently, revised less, and used AI tools dramatically less. The mechanism appears to be that structured journaling offloads the organizational work—so the paper-writing stage can be reserved for synthesis and storytelling.
▸ Whether scaffolding deepens reflection is still an open question
The process data is compelling, but what's happening inside the reflection is harder to measure. Determining whether students are genuinely reflecting more deeply—or simply structuring surface-level thoughts more efficiently—will require larger samples and multiple independent human coders.
▸ Format flexibility widens access without cost
Offering dialogue transcripts as an alternative to written journals accommodates different cognitive styles and removes a barrier for students who find isolated writing difficult. There is no measurable cost to allowing this flexibility.
Future Directions
Test depth more rigorously
Larger samples and multiple independent coders are needed before drawing conclusions about reflection quality. The current qualitative findings are suggestive, not conclusive.
Test AI prompts within writing tools
Explore whether additional AI affordances in Rumi—beyond spellcheck and grammar—help or hurt authentic self-reflection. What kinds of prompts scaffold without replacing?
Compare across student populations
EWMBA, Flex EWMBA, and eMBA students may respond differently to structured scaffolding. Replicating this study across cohorts would clarify what generalizes.
Track longitudinal outcomes
Do students who reflect more deeply in school actually lead differently? Connecting reflection quality to post-graduation leadership behavior remains the long-term question.
Download the full poster
Presented at the Lecturer Teaching Fellows Program showcase, UC Berkeley, 2025–26.
Download PDF