<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://jarnethys.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://jarnethys.com/" rel="alternate" type="text/html" /><updated>2026-07-06T20:28:36+00:00</updated><id>https://jarnethys.com/feed.xml</id><title type="html">Jarne Thys</title><subtitle>Jarne Thys is a Ph.D. researcher at Hasselt University working on human-centered AI for education, focusing on personalized learning for manual skills in structured environments such as school labs, assembly lines, and clinical settings.</subtitle><author><name>Jarne Thys</name><email>jarne.thys -at- uhasselt.be</email></author><entry><title type="html">Scaffolding the Scaffold: Co-Evolving Human-Centered Boundaries for GenAI in Education</title><link href="https://jarnethys.com/ai%20in%20education/human-centered%20ai/adaptive%20ai/generative%20ai/scaffolding-the-scaffold/" rel="alternate" type="text/html" title="Scaffolding the Scaffold: Co-Evolving Human-Centered Boundaries for GenAI in Education" /><published>2026-07-02T00:00:00+00:00</published><updated>2026-07-02T00:00:00+00:00</updated><id>https://jarnethys.com/ai%20in%20education/human-centered%20ai/adaptive%20ai/generative%20ai/scaffolding-the-scaffold</id><content type="html" xml:base="https://jarnethys.com/ai%20in%20education/human-centered%20ai/adaptive%20ai/generative%20ai/scaffolding-the-scaffold/"><![CDATA[<blockquote>
  <table>
    <tbody>
      <tr>
        <td><strong>Related Publication:</strong> <a href="/publication/02/07/2026-hhai-scaffolding-scaffold">Scaffolding the Scaffold: Co-Evolving Human-Centered Boundaries for GenAI in Education</a></td>
        <td><a href="https://doi.org/10.3233/FAIA260504">Full Paper</a></td>
      </tr>
    </tbody>
  </table>
</blockquote>

<p>This blog post describes research presented at HHAI 2026. Co-authored with Yarne Dirkx, Davy Vanacken, and Gustavo Rovelo Ruiz.</p>

<h2 id="what-we-did">What We Did</h2>

<p>Educators today are largely stuck choosing between two extremes for GenAI: ban it, and risk losing real pedagogical opportunities, or allow it, and risk students leaning on it so much that foundational skills like writing, critical thinking, and problem-solving never fully develop. Surveys of top university guidelines show that the vast majority already push instructors toward course-specific GenAI rules, but most of these rules still address only surface-level concerns like plagiarism, and apply the same policy to every student in a course regardless of where they are in their learning journey.</p>

<p>That last part is the gap we focus on. Decades of research on scaffolding tell us that effective support should adapt as a student progresses, staying within their Zone of Proximal Development (ZPD): challenging enough to be useful, but not so far beyond their current ability that it becomes just doing the work for them. A one-size-fits-all GenAI policy can’t do that. A student who has already mastered a skill and a student who has never attempted it are held to the exact same boundary, even though research shows the same tool can help one and hurt the other.</p>

<p>So instead of asking “Can this student use GenAI?”, a binary question that ignores where the student actually is, we ask: <strong>what interactions with GenAI serve this student’s growth right now?</strong> This paper proposes a reconceptualization: adaptive boundaries, GenAI constraints that co-evolve with a student’s demonstrated competency, rather than static rules set once at the course level.</p>

<h2 id="our-framework-adaptive-boundaries-as-pedagogical-scaffolds">Our Framework: Adaptive Boundaries as Pedagogical Scaffolds</h2>

<h3 id="a-three-party-negotiation">A Three-Party Negotiation</h3>

<p>We frame adaptive boundaries as a human-AI collaborative negotiation between three parties:</p>

<ol>
  <li><strong>Students</strong>, who signal competency through their learning behaviors — for example, the depth of questions they ask a GenAI system, patterns in how they complete tasks, or their ability to critically evaluate GenAI output.</li>
  <li><strong>Instructors</strong>, who define learning goals and pedagogical milestones — for example, a writing instructor might require students to demonstrate proper argumentation before GenAI-assisted writing is unlocked, while a programming instructor might allow code generation immediately but require students to explain and modify whatever it produces.</li>
  <li><strong>The GenAI system itself</strong>, which mediates between the two by adapting which of its own features are actually available to a given student at a given time.</li>
</ol>

<p>Instructors set the milestones; the GenAI system then acts as an adaptive layer that scaffolds its own available features based on each student’s detected competency. Critically, this adaptation is bidirectional — boundaries loosen as competency grows, but they can also tighten again during assessments, or if a student’s interaction patterns start to suggest over-reliance rather than growth.</p>

<h3 id="three-developmental-stages">Three Developmental Stages</h3>

<p>Drawing on established models of skill acquisition (the Dreyfus model) and mastery learning (Bloom), we describe three stages where GenAI plays a fundamentally different role:</p>

<ul>
  <li><strong>Novice stage — most restrictive.</strong> Early, unrestricted GenAI access risks undermining exactly the foundational skills a student is trying to build, and the same tool can help one student while hurting another, so the system first needs to build a model of the student before personalizing anything. At this stage, GenAI should expose its reasoning, require metacognitive reflection before answering, and only give direct solutions after a genuine attempt has been made.</li>
  <li><strong>Intermediate stage — more flexible, but monitored.</strong> Boundaries loosen enough to target specific competency gaps, but the main risk shifts to over-reliance: students offloading cognitive work onto GenAI instead of building their own skills. The system should watch for signs like passive help-seeking without articulating the underlying knowledge gap, or uncritical acceptance of AI output, and temporarily tighten access if those patterns show up.</li>
  <li><strong>Advanced stage — largely open.</strong> GenAI functions closer to a professional tool or collaborator. Even here, boundaries tighten again during assessments where students must demonstrate independent competency, and — importantly — boundaries should include aspirational scaffolds that push slightly beyond what’s already been demonstrated, so a competency-gated system doesn’t accidentally cap a student’s growth at what they’ve already shown.</li>
</ul>

<p>Together, these stages turn what looks like an access-control policy into an actual pedagogical instrument that grows with the student.</p>

<h2 id="why-this-is-feasible-now">Why This Is Feasible Now</h2>

<p>We didn’t want this to be purely speculative, so a large part of the paper argues that the pieces needed to build adaptive boundaries already exist — just not combined yet.</p>

<p><strong>On the pedagogical side</strong>, the ZPD and scaffolding theory already establish <em>why</em> support should shrink as competency grows; mastery learning already establishes competency-based (rather than time-based) progression; and Intelligent Tutoring Systems already prove that fine-grained, competency-based adaptation works at scale. Adaptive boundaries extend that logic in two ways: instead of adapting content inside a closed, instructor-authored system, they adapt a student’s <em>access to an external tool</em> used across tasks and courses, and instead of treating the tutoring system as the adaptive instrument, they treat governance itself as the thing that adapts.</p>

<p><strong>On the technical side</strong>, learning analytics research has already shown that interaction data — clickstreams, process mining, language analysis of student-GenAI conversations, and classifiers trained on interaction patterns — can reliably surface competency signals. Meanwhile, mainstream GenAI platforms (ChatGPT, Claude, Ollama) already support the kind of personalized, configurable behavior adaptive boundaries need through system prompts, meaning a single GenAI system could plausibly implement all three stages via prompting rather than requiring bespoke infrastructure per stage.</p>

<p>So the missing piece isn’t a new capability — it’s the conceptual integration of governance and configuration with developmental scaffolding: boundaries that are student-specific, evidence-driven, and designed to evolve as competency changes.</p>

<h2 id="future-research-directions">Future Research Directions</h2>

<p>We lay out four research questions that we think need to be answered before adaptive boundaries can move from concept to reality:</p>

<ul>
  <li><strong>RQ1 — Negotiation:</strong> How should boundaries be negotiated among students, teachers, and GenAI systems, balancing teacher authority, student agency, and equity vs. personalization?</li>
  <li><strong>RQ2 — Detection:</strong> What interaction patterns reliably indicate a student is actually ready for a boundary shift, without oversimplifying (e.g., mistaking a short-term performance spike for real understanding)?</li>
  <li><strong>RQ3 — Transparency:</strong> How can the logic behind a boundary change be made understandable to students, teachers, and administrators alike, without creating cognitive overload or over-explaining?</li>
  <li><strong>RQ4 — Evaluation:</strong> How do we tell whether adaptive boundaries actually work — not just immediate task performance, but durable skill retention, appropriate reliance on GenAI, and student agency — evaluated against both no-GenAI and static-boundary baselines, not just each other?</li>
</ul>

<p>These questions also surface real risks worth naming up front: competency signals may not work equally well across different student backgrounds, “readiness” is hard to standardize across disciplines, and a boundary that scaffolds one student might quietly cap another.</p>

<h2 id="conclusion">Conclusion</h2>

<p>The core argument of this paper is that GenAI governance in education doesn’t have to be a binary policy decision applied the same way to everyone. The pedagogical theory (ZPD, scaffolding, mastery learning) and the technical capability (learning analytics, configurable GenAI) to do better already exist — what’s missing is putting them together into boundaries that are student-specific, evidence-driven, and built to evolve.</p>

<p>If we extend this work further, the most important next steps are the four research questions themselves: running participatory sessions to understand how negotiation should actually work, testing which interaction signals genuinely predict readiness, prototyping transparency mechanisms that scale from simple to in-depth explanations, and designing the mixed-methods, multi-baseline studies needed to evaluate whether any of this actually improves learning outcomes.</p>

<p>Ultimately, we want GenAI in education to stop being a fixed policy line and start being what scaffolding always was supposed to be: something that knows when to help, when to challenge, and when to step back.</p>

<hr />

<p><em>This blog post is based on research presented at HHAI 2026. For complete details, see the <a href="https://doi.org/10.3233/FAIA260504">full paper</a>.</em></p>]]></content><author><name>Jarne Thys</name><email>jarne.thys -at- uhasselt.be</email></author><category term="AI in Education" /><category term="Human-Centered AI" /><category term="Adaptive AI" /><category term="Generative AI" /><summary type="html"><![CDATA[A vision for GenAI access in education that adapts per student, tightening or loosening as competency develops, grounded in Zone of Proximal Development and scaffolding theory.]]></summary></entry><entry><title type="html">Engineering Trustworthy Automation: Design Principles for AutoML Tools for Novices</title><link href="https://jarnethys.com/automl/human-computer%20interaction/machine%20learning/ux%20research/trustworthy-automl-for-novices/" rel="alternate" type="text/html" title="Engineering Trustworthy Automation: Design Principles for AutoML Tools for Novices" /><published>2025-11-27T00:00:00+00:00</published><updated>2025-11-27T00:00:00+00:00</updated><id>https://jarnethys.com/automl/human-computer%20interaction/machine%20learning/ux%20research/trustworthy-automl-for-novices</id><content type="html" xml:base="https://jarnethys.com/automl/human-computer%20interaction/machine%20learning/ux%20research/trustworthy-automl-for-novices/"><![CDATA[<blockquote>
  <table>
    <tbody>
      <tr>
        <td><strong>Related Publication:</strong> <a href="/publication/27/11/2025-automl-for-novices">Engineering Trustworthy Automation: Design Principles and Evaluation for AutoML Tools for Novices</a></td>
        <td><a href="https://arxiv.org/abs/2511.22352">arXiv Paper</a></td>
      </tr>
    </tbody>
  </table>
</blockquote>

<p>This blog post describes research presented at the EICS 2025 International Workshops and Doctoral Consortium.</p>

<h2 id="what-we-did">What We Did</h2>

<p>This work focused on a practical problem: many people want to use machine learning for their own domain tasks, but most AutoML tools are still not truly designed for beginners. Existing systems often automate model search and optimization, yet they leave major gaps in usability, understanding, and end-to-end workflow support. As a result, novices may be able to click through a pipeline, but still not understand what the system is doing, when to trust it, or how to use the trained model safely afterward.</p>

<p>The goal of this research was to design and evaluate a more trustworthy form of automation for novice users. Instead of focusing on peak algorithmic performance, we focused on building an AutoML workflow that helps users successfully create a working model, understand the main steps, and maintain a sense of control throughout the process.</p>

<p>More specifically, the project had three core contributions:</p>

<ol>
  <li>
    <p><strong>An abstract AutoML pipeline for novices</strong>
We proposed an end-to-end pipeline covering the full process from data intake to inference, rather than only automating isolated parts of model training.</p>
  </li>
  <li>
    <p><strong>A prototype implementation called NovaClass</strong>
To examine that pipeline in practice, we built a novice-oriented system for Transformer-based text classification that makes advanced model training more accessible through a guided interface.</p>
  </li>
  <li>
    <p><strong>A user study and resulting design principles</strong>
We evaluated the prototype with 24 participants and used the results to derive four design principles for future AutoML tools aimed at novices: first-model success, explanations, context-aware abstractions, and predictability through safeguards.</p>
  </li>
</ol>

<p>This problem matters because domain experts increasingly want to apply modern AI to their own work—whether that means grading open-ended answers, analyzing text, or building task-specific classifiers—but they are often blocked by programming requirements, framework complexity, and fragile configuration steps. If AutoML tools are going to make AI genuinely more accessible, they need to support not just automation, but also trust, usability, and learning.</p>

<h2 id="how-we-did-it">How We Did It</h2>

<h3 id="1-designing-an-abstract-pipeline">1. Designing an Abstract Pipeline</h3>

<p>We first designed an abstract end-to-end AutoML pipeline specifically aimed at novice users. The paper frames this pipeline as six connected stages:</p>

<ol>
  <li>
    <p><strong>Data Intake and Upfront Safeguards</strong>
The system begins by ingesting data and validating it early. This includes identifying issues like missing values, checking labels, and catching common problems before training starts. The idea is to reduce avoidable failures and make the system more robust from the start.</p>
  </li>
  <li>
    <p><strong>Key Model Parameter Configuration</strong>
Rather than exposing every technical setting, the system only shows the parameters that matter most for novice decision-making—such as selecting which columns should be used as input and which column should be predicted. Other settings can be handled automatically using safe defaults.</p>
  </li>
  <li>
    <p><strong>One-Click Training</strong>
Once the essential choices are made, the rest of the pipeline is automated. This includes preprocessing, train/validation/test splitting, label conversion, and other steps that usually require technical knowledge.</p>
  </li>
  <li>
    <p><strong>Simplified Results</strong>
After training, the system presents performance in a simplified way so users can understand outcomes without needing deep ML expertise. The emphasis is on interpretability and clarity rather than only raw metrics.</p>
  </li>
  <li>
    <p><strong>Save Model and Metadata</strong>
In addition to saving the trained model, the system stores metadata about the training configuration. This creates an inference “contract” that helps the system remember how the model expects data later.</p>
  </li>
  <li>
    <p><strong>Auto-Configured Inference</strong>
Finally, the saved metadata is used to automatically configure the prediction interface so that users can apply the model consistently without manually rebuilding the setup.</p>
  </li>
</ol>

<p>This pipeline was intentionally designed around <strong>reliability and novice comprehension</strong>, not maximum benchmark performance. The core idea is that a good beginner tool should make it easy to reach a safe and usable baseline first.</p>

<h3 id="2-building-a-prototype-novaclass">2. Building a Prototype: NovaClass</h3>

<p>To put the abstract pipeline into practice, we implemented <strong>NovaClass</strong>, a prototype system for novice-friendly automation of <strong>Transformer-based text classification</strong>. We chose text classification because it is widely useful—for example in grading, spam detection, and sentiment or emotion analysis—while still being complex enough that current Transformer workflows are difficult for non-experts to configure manually.</p>

<p>NovaClass includes several key features:</p>

<ul>
  <li>
    <p><strong>Dataset inspection at upload time</strong>
Users upload a CSV file, and the system highlights column types, missing rows, label balance, and basic statistics. Users can also inspect class distributions and preview the first rows of the dataset. This helps users understand their data before training begins.</p>
  </li>
  <li>
    <p><strong>A simplified configuration interface</strong>
The interface only exposes decisions novices are expected to understand, especially which columns contain the input text and which column contains the target labels. On page 6 of the paper, the interface is shown with configuration options on the left and an integrated contextual assistant on the right, designed to explain choices in plain language.</p>
  </li>
  <li>
    <p><strong>Safe defaults and automated training</strong>
The system runs a reproducible pipeline with sensible defaults, aiming to help users produce a working classifier on their first attempt. Metadata is saved automatically to reduce mismatches later during inference.</p>
  </li>
  <li>
    <p><strong>One-toggle cascade classification</strong>
NovaClass also includes a simplified way to use a more advanced strategy: cascade classification. Instead of training a single multi-class model, the system can automatically decompose the task into a hierarchy of binary classifiers. On page 7, the paper shows how this is presented in the interface as a one-toggle option, making a more advanced technique accessible without additional configuration burden.</p>
  </li>
  <li>
    <p><strong>Metadata-driven inference</strong>
The inference interface uses saved metadata such as label order, encoders, and strategy so that users can apply the model without manually reconstructing the training setup. The system also shows both the predicted label confidence and full class-probability distribution.</p>
  </li>
  <li>
    <p><strong>A conversational assistant</strong>
NovaClass includes a context-aware assistant that explains metrics like accuracy, recall, and F1-score, and suggests next steps based on what the user is currently doing. The prototype uses <strong>IBM Granite 3.3 8B</strong> because it followed instructions well in testing, has a large context window, and is small enough to run locally for privacy-sensitive use cases.</p>
  </li>
</ul>

<h3 id="3-conducting-a-user-study">3. Conducting a User Study</h3>

<p>We evaluated NovaClass through a <strong>24-participant study</strong> designed to assess usability, trust, and understanding. The study included users with different levels of prior ML experience and compared novices against participants who had previously trained models.</p>

<p>Participants completed three tasks of increasing complexity:</p>

<ol>
  <li>
    <p><strong>Task 1: Binary text classification</strong>
Participants configured and trained a classifier for fake news detection. This tested whether the guided configuration helped them choose the right input and label columns.</p>
  </li>
  <li>
    <p><strong>Task 2: Cascade classification</strong>
Participants trained a cascaded classifier for e-commerce description categorization. This tested whether the interface could make a more advanced classification strategy approachable.</p>
  </li>
  <li>
    <p><strong>Task 3: Diagnosis task</strong>
Participants analyzed results to identify a class imbalance issue in a dataset. This tested whether the available tools helped users interpret model performance and diagnose problems.</p>
  </li>
</ol>

<p>To measure outcomes, we collected:</p>

<ul>
  <li>demographic and background information,</li>
  <li><strong>PAILQ-6</strong> scores for perceived AI literacy,</li>
  <li>a questionnaire on <strong>trust and understandability</strong> based on prior AutoML trust research,</li>
  <li>and the <strong>User Experience Questionnaire (UEQ)</strong>.</li>
</ul>

<h2 id="results">Results</h2>

<p>The study showed that the system worked well as a usable end-to-end workflow, but it also revealed an important gap between task completion and genuine understanding.</p>

<h3 id="success-metrics">Success Metrics</h3>

<p>The strongest result was that <strong>all 24 participants successfully trained a working binary classifier and a working cascaded classifier</strong>. That means the core workflow succeeded in making advanced Transformer-based classification accessible enough for everyone in the study to complete the main tasks.</p>

<p>More detailed task outcomes were:</p>

<ul>
  <li><strong>Task 1</strong>: All participants trained a functioning fake news classifier. Also, <strong>21 of 24 participants (87.5%)</strong> reported at least some confidence in training the binary model afterward.</li>
  <li><strong>Task 2</strong>: All participants correctly trained a cascade model, and <strong>22 of 24 participants (91.7%)</strong> correctly identified the weakest-performing stage in the cascade.</li>
  <li><strong>Task 3</strong>: <strong>17 of 24 participants (70.8%)</strong> correctly identified label imbalance as the main issue in the dataset.</li>
</ul>

<p>These results suggest that the combination of guided configuration, automation, and metadata-driven inference successfully reduced common setup errors that would normally block beginners.</p>

<h3 id="user-experience">User Experience</h3>

<p>Participants rated the system positively across all six UEQ dimensions. The highest scores were:</p>

<ul>
  <li><strong>Efficiency</strong>: 1.917</li>
  <li><strong>Attractiveness</strong>: 1.778</li>
  <li><strong>Perspicuity</strong>: 1.635</li>
</ul>

<p>The remaining dimensions were also positive:</p>

<ul>
  <li><strong>Dependability</strong>: 1.375</li>
  <li><strong>Stimulation</strong>: 1.458</li>
  <li><strong>Novelty</strong>: 1.052</li>
</ul>

<p>The figure on page 10 shows these ratings against the UEQ benchmark, where the strongest results are in efficiency, attractiveness, and perspicuity, with more moderate but still positive scores for dependability and stimulation.</p>

<p>So from a pure UX perspective, NovaClass was received as effective, appealing, and relatively easy to understand.</p>

<h3 id="trust-and-understanding">Trust and Understanding</h3>

<p>The more nuanced finding was that <strong>experienced users reported significantly higher trust and understanding than novices</strong>. This was not just a minor trend; the difference between the groups was statistically significant for the overall trust/understandability score. The median average rating was <strong>4.05</strong> for experienced users versus <strong>3.64</strong> for novices.</p>

<p>Significant differences also appeared on individual questions, including:</p>

<ul>
  <li>understanding the tool,</li>
  <li>understanding the overall process,</li>
  <li>understanding the data,</li>
  <li>and understanding model evaluation metrics.</li>
</ul>

<p>The biggest gap appeared in understanding evaluation metrics: experienced users had a median rating of <strong>5</strong>, while novices had a median of <strong>2.5</strong>. This is especially important because metrics are central to deciding whether a model is actually good enough to use.</p>

<h3 id="deployment-confidence">Deployment Confidence</h3>

<p>When asked whether they would deploy models trained with NovaClass, <strong>17 of 24 participants (70.8%)</strong> said yes. Most justified this with statements about “high accuracy” or “high F1-score.” Meanwhile, the participants who said no often cited transparency concerns, such as not knowing what the model based its decisions on or not understanding how the data had been processed.</p>

<p>This result shows a key tension:</p>

<ul>
  <li>some users may be too willing to trust strong headline metrics,</li>
  <li>while others hesitate because they lack insight into the model’s reasoning and the system’s hidden steps.</li>
</ul>

<p>In other words, automation can help users finish tasks, but without better explanations, it does not automatically produce appropriate reliance.</p>

<h3 id="role-of-the-conversational-assistant">Role of the Conversational Assistant</h3>

<p>The conversational assistant played a major role across all tasks. Participants frequently used it:</p>

<ul>
  <li>to select inputs in Task 1,</li>
  <li>to identify the weak stage in Task 2,</li>
  <li>and to diagnose imbalance in Task 3.</li>
</ul>

<p>In some cases, it was used as often as or more often than static tools like the classification report or confusion matrix. This suggests that context-aware, plain-language guidance is especially valuable for novice users trying to make sense of ML workflows. At the same time, the paper notes that this also raises the cost of occasional hallucinations or imprecise explanations.</p>

<h3 id="four-design-principles">Four Design Principles</h3>

<p>Based on the study and relevant theory, we derived four design principles for future AutoML systems targeting novices.</p>

<p><strong>P1: Support First-Model Success</strong>
The system should make it easy for users to produce a working baseline on their first attempt. Early success raises self-efficacy, builds momentum, and reduces the chance that beginners quit before they see value. In practice, this means safe defaults, aggressive upfront validation, automatic preprocessing, one-click training, and advanced functionality that can still be enabled with minimal friction.</p>

<p><strong>P2: Provide Explanations to Build Mental Models and Appropriate Reliance</strong>
Metrics alone are not enough. Users need simplified, interpretable explanations of what a score means, how to read a visualization, and what limitations remain. The goal is not just to inform, but to help users avoid both over-trust and under-trust. The study’s deployment results showed exactly why this matters.</p>

<p><strong>P3: Use Abstractions and Context-Aware Assistance to Keep Users in Their Zone of Proximal Development</strong>
Interfaces should hide unnecessary complexity, but still provide targeted support that helps users perform just beyond what they could do alone. This includes adaptive interfaces, stage-aware assistance, and advanced features that stay easy to access. NovaClass’s assistant is one example of this principle in practice.</p>

<p><strong>P4: Ensure Predictability and Safeguards to Strengthen Perceived Control</strong>
Trust also depends on consistency and control. Systems should be predictable, validate data and settings before training, generate inference interfaces directly from saved metadata, and prevent train-inference mismatches. These safeguards make the workflow feel safer and reduce invisible failure points.</p>

<h2 id="conclusion">Conclusion</h2>

<p>This work shows that making AutoML accessible to novices is not just a technical automation problem. It is equally a <strong>human-computer interaction problem</strong>. If a tool only automates the hard parts without helping users understand what is happening, then it may reduce effort, but it will not necessarily build trust, correct mental models, or appropriate confidence.</p>

<p>NovaClass demonstrated that a carefully designed workflow can make advanced tasks like Transformer fine-tuning and cascade classification usable even for beginners: every participant completed the main training tasks, and overall UX ratings were positive. But the study also made clear that successful task completion does not automatically mean deep understanding. Experienced users still understood the workflow, the data, and the metrics better than novices.</p>

<p>The main lesson, then, is that good AutoML design for novices should balance:</p>

<ul>
  <li><strong>automation</strong> so people can get started,</li>
  <li><strong>explanations</strong> so they understand what happened,</li>
  <li><strong>assistance</strong> so they can keep progressing,</li>
  <li>and <strong>safeguards</strong> so they remain in control.</li>
</ul>

<p>If we extend this work further, the most important next steps would be:</p>

<ul>
  <li>building <strong>expertise-adaptive interfaces</strong> that reveal more detail as users grow,</li>
  <li>improving <strong>explainability</strong> so metrics and model behavior are easier to interpret,</li>
  <li>and studying how these systems affect learning and confidence over time through <strong>longitudinal evaluations</strong>.</li>
</ul>

<p>Ultimately, the project argues for a more human-centered vision of AutoML: not tools that simply automate model building, but tools that help people become more capable, more informed, and more appropriately confident while using AI.</p>

<hr />

<p><em>This blog post is based on research presented at EICS 2025. For complete technical details and evaluation results, see the <a href="https://arxiv.org/abs/2511.22352">full paper on arXiv</a>.</em></p>]]></content><author><name>Jarne Thys</name><email>jarne.thys -at- uhasselt.be</email></author><category term="AutoML" /><category term="Human-Computer Interaction" /><category term="Machine Learning" /><category term="UX Research" /><summary type="html"><![CDATA[How NovaClass and a 24-participant study produced four design principles for more understandable, predictable, and usable AutoML tools for novices.]]></summary></entry><entry><title type="html">The Effect of VR and Haptic Feedback Devices on Online Consumer Decision-Making</title><link href="https://jarnethys.com/virtual%20reality/human-computer%20interaction/e-commerce/haptic%20feedback/vr-haptic-feedback-ecommerce/" rel="alternate" type="text/html" title="The Effect of VR and Haptic Feedback Devices on Online Consumer Decision-Making" /><published>2025-07-07T00:00:00+00:00</published><updated>2025-07-07T00:00:00+00:00</updated><id>https://jarnethys.com/virtual%20reality/human-computer%20interaction/e-commerce/haptic%20feedback/vr-haptic-feedback-ecommerce</id><content type="html" xml:base="https://jarnethys.com/virtual%20reality/human-computer%20interaction/e-commerce/haptic%20feedback/vr-haptic-feedback-ecommerce/"><![CDATA[<blockquote>
  <table>
    <tbody>
      <tr>
        <td><strong>Related Publication:</strong> <a href="/publication/07/07/2025-rarcs-abstract">The Effect of VR and Haptic Feedback Devices on Online Consumer Decision-Making</a></td>
        <td><a href="http://hdl.handle.net/1942/48289">Conference Abstract</a></td>
      </tr>
    </tbody>
  </table>
</blockquote>

<p>This blog post describes research presented at the 31st Recent Advances in Retailing &amp; Consumer Science (RARCS) Conference. Co-authored with Lieve Douce, Stephanie Van De Sanden, Kim Willems, Davy Vanacken, and Gustavo Rovelo Ruiz.</p>

<h2 id="what-we-did">What We Did</h2>

<p>This research examined whether Virtual Reality (VR) and haptic feedback can help consumers make better-informed decisions when evaluating products online, especially for products whose qualities are difficult to judge before purchase.</p>

<p>The core problem is that many products are <strong>experience products</strong>: consumers need to interact with them to properly assess them. In the pre-purchase stage, this is often difficult or impossible. That is especially true when decision-making happens online, where sensory information is limited. In our work, we focused on windows as a concrete example. When building or renovating a house, consumers may struggle to anticipate:</p>

<ul>
  <li>how smoothly and ergonomically a sliding window opens, and</li>
  <li>whether investing in more insulating glazing will improve indoor comfort enough to justify the extra cost.</li>
</ul>

<p>The research question was whether combining richer visual information through VR with tactile input through haptic feedback could reduce this uncertainty and improve evaluation, confidence, and satisfaction.</p>

<p>This matters for e-commerce and digital product experiences because online interfaces are often visually rich but sensorily poor. From an HCI perspective, the study explores how multisensory interfaces can support better user judgment by helping people form more vivid and realistic mental representations of products before they buy.</p>

<h2 id="how-we-did-it">How We Did It</h2>

<p>We conducted <strong>two experimental lab studies</strong> on consumer decision-making about windows. The studies were grounded in <strong>mental imagery theory</strong>, which suggests that richer sensory input helps consumers form more vivid mental images of a product, making it easier to understand how it would function in real life.</p>

<h3 id="research-design">Research Design</h3>

<p>Both experiments used a <strong>between-subjects 2×2 design</strong>:</p>

<ol>
  <li>
    <p><strong>Visual Information</strong></p>

    <ul>
      <li>Video</li>
      <li>VR</li>
    </ul>
  </li>
  <li>
    <p><strong>Haptic Information</strong></p>

    <ul>
      <li>No haptic feedback</li>
      <li>Simulated haptic feedback</li>
    </ul>
  </li>
</ol>

<p>This created four conditions in each experiment, allowing us to compare the separate and combined effects of immersive visualization and touch-based feedback.</p>

<p>The two use cases were:</p>

<ol>
  <li>
    <p><strong>Sliding window system</strong></p>

    <ul>
      <li>Participants evaluated different window sliding mechanisms.</li>
      <li>The haptic feedback was delivered through a <strong>mechanical force-feedback instrument</strong> built to simulate the amount of force needed to open a sliding window.</li>
    </ul>
  </li>
  <li>
    <p><strong>Window glazing insulation</strong></p>

    <ul>
      <li>Participants evaluated different glazing options, specifically <strong>double vs. triple glazing</strong>.</li>
      <li>The haptic feedback was delivered through an <strong>off-the-shelf haptic glove</strong>, which simulated differences in insulation value.</li>
    </ul>
  </li>
</ol>

<h3 id="study-protocol">Study Protocol</h3>

<p>Participants were placed in a scenario where they had to evaluate <strong>two alternatives</strong>:</p>

<ul>
  <li>either two types of window sliding systems, or</li>
  <li>two types of window glazing.</li>
</ul>

<p>The total sample consisted of <strong>159 consumers in Belgium</strong>:</p>

<ul>
  <li><strong>Average age:</strong> 39.77 years</li>
  <li><strong>SD:</strong> 12.34</li>
  <li><strong>Female participants:</strong> 100 out of 159 (62.9%)</li>
</ul>

<p>After interacting with the assigned condition, participants completed measures related to both the medium and the product.</p>

<p>We collected data on:</p>

<ul>
  <li>mental imagery</li>
  <li>ease of evaluation</li>
  <li>perceived informativeness</li>
  <li>multiple customer experience dimensions</li>
  <li>processing fluency</li>
  <li>satisfaction</li>
  <li>preference</li>
</ul>

<p>In addition, <strong>eye-tracking data</strong> were logged during the experiments.</p>

<p>By comparing video versus VR and no haptics versus simulated haptics, we could isolate the contribution of each modality and their combined effect on consumer decision-making.</p>

<h3 id="technical-implementation">Technical Implementation</h3>

<p>The technical setup included:</p>

<ul>
  <li><strong>Visual platform:</strong> a comparison between standard video-based presentation and immersive VR</li>
  <li>
    <p><strong>Haptic devices:</strong></p>

    <ul>
      <li>a custom mechanical force-feedback device for the sliding-window scenario</li>
      <li>an off-the-shelf haptic glove for the glazing-insulation scenario</li>
    </ul>
  </li>
  <li><strong>Virtual environment:</strong> product evaluation scenarios designed to help participants compare realistic window alternatives</li>
  <li><strong>Data collection:</strong> self-report survey measures combined with eye-tracking logs</li>
</ul>

<p>Rather than simulating a generic virtual store, the system was tailored to realistic home-building and renovation decisions, where consumers must judge functional product characteristics that are hard to assess from standard online content alone.</p>

<h2 id="results">Results</h2>

<p>The results showed clear positive effects of enhanced visual and haptic input on consumer decision-making.</p>

<h3 id="main-findings">Main Findings</h3>

<ol>
  <li>
    <p><strong>Mental Imagery and Understanding</strong></p>

    <ul>
      <li>VR and haptic feedback improved participants’ ability to form vivid mental images of the product.</li>
      <li>Richer sensory input made it easier for participants to imagine how the product would work in real life.</li>
    </ul>
  </li>
  <li>
    <p><strong>Ease of Evaluation and Informativeness</strong></p>

    <ul>
      <li>Participants found the products easier to evaluate when more sensory information was available.</li>
      <li>VR and haptic feedback increased the perceived informativeness of the experience.</li>
    </ul>
  </li>
  <li>
    <p><strong>Customer Experience and Processing Fluency</strong></p>

    <ul>
      <li>Enhanced visual and tactile input improved overall customer experience.</li>
      <li>Participants processed the information more fluently, suggesting that multisensory interaction made the decision task feel more natural and intuitive.</li>
    </ul>
  </li>
  <li>
    <p><strong>Satisfaction</strong></p>

    <ul>
      <li>Consumers reported greater satisfaction with their decision-making experience when VR and haptic feedback were included.</li>
    </ul>
  </li>
  <li>
    <p><strong>Higher-Investment Choices</strong></p>

    <ul>
      <li>For more expensive or higher-stakes options—such as <strong>triple-glazed windows</strong> or <strong>advanced sliding systems</strong>—haptic feedback played a particularly important role.</li>
      <li>In these cases, touch-based simulation was more influential in persuading consumers to invest in options that improve comfort, such as smoother sliding mechanisms or better indoor climate performance.</li>
    </ul>
  </li>
</ol>

<h3 id="implications-for-e-commerce">Implications for E-Commerce</h3>

<p>The findings suggest that VR and haptic feedback can:</p>

<ul>
  <li>provide richer product information before purchase</li>
  <li>improve consumers’ ability to evaluate experience products online</li>
  <li>reduce uncertainty in higher-stakes purchase decisions</li>
  <li>increase confidence and satisfaction in digital buying journeys</li>
</ul>

<p>These technologies appear especially valuable for products where functionality, comfort, or physical interaction matter more than appearance alone. In such cases, adding haptic feedback may be more impactful than visual enhancement alone.</p>

<p>For e-commerce, this points to a future where multisensory interfaces are not just novel experiences, but practical tools for helping consumers make better decisions.</p>

<h2 id="conclusion">Conclusion</h2>

<p>This research demonstrates that combining VR and haptic feedback can meaningfully improve online consumer decision-making by giving consumers richer sensory information before purchase.</p>

<p><strong>Key Takeaways</strong>:</p>

<ol>
  <li>VR and haptic feedback improve mental imagery, ease of evaluation, perceived informativeness, processing fluency, customer experience, and satisfaction.</li>
  <li>Haptic feedback is especially important for higher-investment or comfort-related product decisions.</li>
  <li>Multisensory interaction can help bridge the gap between online and offline product evaluation.</li>
</ol>

<p>From an HCI perspective, the work reinforces the importance of designing systems that go beyond visual presentation alone. When people can better simulate interaction with a product, they can make more informed, confident, and satisfying choices.</p>

<p><strong>Future Directions</strong>:</p>

<ul>
  <li>Investigating long-term adoption of VR shopping</li>
  <li>Exploring more sophisticated haptic feedback (temperature, texture variation)</li>
  <li>Studying different product categories and consumer segments</li>
  <li>Examining the role of social presence in VR shopping environments</li>
  <li>Cost-effectiveness analysis for e-commerce businesses</li>
</ul>

<hr />

<p><em>This blog post is based on research presented at the 31st RARCS Conference, 2025. For details, see the <a href="http://hdl.handle.net/1942/48289">conference abstract</a>.</em></p>]]></content><author><name>Jarne Thys</name><email>jarne.thys -at- uhasselt.be</email></author><category term="Virtual Reality" /><category term="Human-Computer Interaction" /><category term="E-Commerce" /><category term="Haptic Feedback" /><summary type="html"><![CDATA[How two laboratory studies tested whether virtual reality and haptic feedback improve online product evaluation, confidence, and decision-making.]]></summary></entry><entry><title type="html">Improving AI Text Classification: A Cascaded Approach for Educational Grading</title><link href="https://jarnethys.com/ai%20in%20education/machine%20learning/text%20classification/cascaded-ai-text-classification/" rel="alternate" type="text/html" title="Improving AI Text Classification: A Cascaded Approach for Educational Grading" /><published>2025-06-24T00:00:00+00:00</published><updated>2025-06-24T00:00:00+00:00</updated><id>https://jarnethys.com/ai%20in%20education/machine%20learning/text%20classification/cascaded-ai-text-classification</id><content type="html" xml:base="https://jarnethys.com/ai%20in%20education/machine%20learning/text%20classification/cascaded-ai-text-classification/"><![CDATA[<blockquote>
  <table>
    <tbody>
      <tr>
        <td><strong>Related Publication:</strong> <a href="/publication/01/07/2025-ai-text-cascaded">Improving AI Text Classification: A Cascaded Approach</a></td>
        <td><a href="http://hdl.handle.net/1942/46328">Full Paper</a></td>
      </tr>
    </tbody>
  </table>
</blockquote>

<p>This blog post describes research presented at the 3rd Workshop on Engineering Interactive Systems Embedding AI Technologies in Trier, Germany (June 2025).</p>

<h2 id="what-we-did">What We Did</h2>

<p>This work focused on improving the reliability of AI-assisted grading for open-ended student answers. While Large Language Models (LLMs) are flexible and easy to repurpose, they still struggle with consistent and trustworthy classification in high-stakes educational tasks such as grading. In particular, they tend to produce fluent output while still making classification mistakes, and they often avoid assigning low grades such as “incorrect.”</p>

<p>The core goal was to engineer a grading pipeline that is both more reliable and more usable in practice. We wanted a system that could:</p>

<ul>
  <li>classify short answers at the rubric level more accurately,</li>
  <li>improve detection of incorrect answers,</li>
  <li>preserve the flexibility of LLMs for feedback generation,</li>
  <li>and keep teachers in control through a transparent review interface.</li>
</ul>

<p>The specific grading setup used three labels: <strong>Incorrect</strong>, <strong>Partially correct</strong>, and <strong>Correct</strong>. Rather than asking a single model to both decide the grade and generate feedback, we separated these tasks. The classification step is handled first by dedicated transformer models, and only then is the result passed to an LLM-based feedback layer. This makes the feedback generation more grounded, because the LLM no longer has to infer the grade by itself.</p>

<p>A major motivation behind this work was that existing approaches were not sufficient for dependable educational use. A generic LLM baseline (Gemma 3 27B) achieved only <strong>58% accuracy</strong> on the evaluation set and showed a clear bias toward marking answers as correct, identifying only <strong>21%</strong> of incorrect answers. That kind of failure mode is especially problematic in grading, where missing incorrect answers undermines both fairness and usefulness.</p>

<h2 id="how-we-did-it">How We Did It</h2>

<p>We designed an end-to-end approach with two main parts: a <strong>transformer cascade for rubric-level classification</strong> and a <strong>Mixture-of-Agents (MoA) LLM system</strong> for generating feedback.</p>

<h3 id="1-transformer-cascade-for-classification">1. Transformer Cascade for Classification</h3>

<p>Instead of training a single model to directly predict all three grading labels at once, we implemented a cascaded architecture using fine-tuned <strong>BERT-base-uncased</strong> models.</p>

<p>The cascade works in two stages:</p>

<ul>
  <li>The first transformer performs a binary classification: <strong>Correct</strong> vs. <strong>Not correct</strong></li>
  <li>If an answer is classified as <strong>Not correct</strong>, it is passed to a second transformer</li>
  <li>The second transformer then distinguishes between <strong>Incorrect</strong> and <strong>Partially correct</strong></li>
</ul>

<p>This decomposition turns one harder three-class problem into two simpler decisions. That matters because the grading labels are imbalanced and adjacent categories share overlapping features. In the single-model setup, the classifier often confused <strong>Correct</strong> and <strong>Partially correct</strong> and was hesitant to label answers as <strong>Incorrect</strong>. By isolating clearly correct answers first, the cascade can focus the second stage on a more balanced subset, making it easier to distinguish the weaker responses.</p>

<p>For input, we supplied both:</p>

<ul>
  <li>the <strong>student answer</strong></li>
  <li>and a <strong>reference answer</strong></li>
</ul>

<p>Although we aimed to keep required metadata minimal, adding the reference answer produced a substantial performance boost. That led to an important practical conclusion: the extra effort of creating reference answers is justified if it significantly improves grading accuracy.</p>

<h3 id="2-mixture-of-agents-llm-system">2. Mixture-of-Agents LLM System</h3>

<p>After classification, the assigned label and the original student input are combined into a prompt for a <strong>Mixture-of-Agents</strong> feedback system.</p>

<p>The purpose of this second layer is not to decide the grade, but to generate clearer and more appropriate feedback based on the already-determined classification. This separation improves alignment between the grade and the explanation. Instead of relying on an LLM to both judge and justify at once, the MoA setup uses the classification result as structured guidance for the feedback stage.</p>

<p>To make this usable for teachers, we paired the feedback mechanism with a <strong>traffic-light interface</strong>:</p>

<ul>
  <li><strong>Green</strong> for fully correct</li>
  <li><strong>Yellow</strong> for partially correct</li>
  <li><strong>Red</strong> for incorrect</li>
</ul>

<p>In the prototype, this idea is applied to programming assignments line by line. Each code line is highlighted with one of the traffic-light colors, and the system generates comments that include reinforcement for correct parts and constructive suggestions for lines that need improvement. The teacher can then accept, revise, or override that feedback.</p>

<h3 id="evaluation">Evaluation</h3>

<p>We evaluated the approach on the <strong>ASAG (Automatic Short Answer Grading)</strong> dataset, using <strong>130 student responses</strong> labeled into the three rubric categories. We compared three systems:</p>

<ul>
  <li>a generic LLM baseline (<strong>Gemma 3 27B</strong>),</li>
  <li>a <strong>single transformer</strong> trained for direct three-class classification,</li>
  <li>and the proposed <strong>cascade model</strong>.</li>
</ul>

<p>The LLM baseline was included as a realistic comparison point because it represents the kind of flexible model many people might want to use directly for grading. The single transformer served as the strongest non-cascaded baseline using the same encoder family. Performance was evaluated with standard classification metrics: <strong>precision</strong>, <strong>recall</strong>, <strong>F1-score</strong>, and overall <strong>accuracy</strong>.</p>

<h2 id="results">Results</h2>

<p>The cascaded approach produced a clear improvement over both baselines.</p>

<p>Compared to the <strong>single transformer</strong>, the cascade:</p>

<ul>
  <li>increased <strong>recall for Incorrect answers</strong> from <strong>0.26 to 0.58</strong></li>
  <li>increased <strong>precision for Correct answers</strong> from <strong>0.70 to 0.84</strong></li>
  <li>improved overall <strong>accuracy</strong> from <strong>0.65 to 0.82</strong></li>
</ul>

<p>This means the cascade more than doubled the system’s ability to correctly identify incorrect answers, while also becoming more precise when marking answers as fully correct. That is exactly the kind of tradeoff improvement that matters in grading: stronger detection of weak answers without becoming overly harsh on strong ones.</p>

<p>The full results for the cascade were:</p>

<ul>
  <li><strong>Incorrect</strong>: Precision 0.85, Recall 0.58, F1 0.69</li>
  <li><strong>Partially Correct</strong>: Precision 0.77, Recall 0.77, F1 0.77</li>
  <li><strong>Correct</strong>: Precision 0.84, Recall 0.91, F1 0.87</li>
  <li><strong>Overall accuracy</strong>: 0.82</li>
</ul>

<p>By contrast:</p>

<ul>
  <li>the <strong>LLM baseline</strong> achieved only <strong>0.58 accuracy</strong> and showed a strong bias toward predicting <strong>Correct</strong>,</li>
  <li>the <strong>single transformer</strong> improved on that with <strong>0.65 accuracy</strong>, but still struggled with confusion between adjacent labels.</li>
</ul>

<p>The confusion matrices in the paper illustrate this clearly: the LLM rarely labeled answers as incorrect, the single transformer still made many off-diagonal errors between <strong>Correct</strong> and <strong>Partially correct</strong>, and the cascade produced a much cleaner diagonal with fewer misclassifications.</p>

<h3 id="prototype-implementation">Prototype Implementation</h3>

<p>Beyond the classification experiment, we also developed a prototype for <strong>semi-automatic grading</strong>.</p>

<p>The prototype was designed to fit into existing teaching workflows with minimal disruption:</p>

<ul>
  <li>it visually marks answer quality using a traffic-light metaphor,</li>
  <li>it generates first-draft feedback for the teaching staff,</li>
  <li>and it keeps a <strong>human-in-the-loop</strong> so staff can review and revise both the grading and the comments before finalizing them.</li>
</ul>

<p>For programming tasks, the interface analyzes code line by line. This gives teachers a more granular overview than grading an entire submission as a single block. The color-coding helps direct attention to the most important problem areas, while the auto-generated comments reduce the amount of feedback that must be written from scratch. The result is a workflow that aims to save time without hiding the reasoning process from the teacher.</p>

<h2 id="conclusion">Conclusion</h2>

<p>This work shows that better AI grading is not just about using a larger model. Careful system design matters. By splitting rubric prediction into a <strong>cascaded transformer architecture</strong> and separating that from an <strong>LLM-based feedback layer</strong>, we were able to build a pipeline that is more accurate, more transparent, and better suited to educational use than a direct one-model approach.</p>

<p>Three lessons stood out most clearly:</p>

<ol>
  <li>
    <p><strong>Cascaded architectures improve reliability</strong>
Breaking the grading task into simpler decisions reduced confusion between adjacent classes and significantly improved detection of incorrect answers.</p>
  </li>
  <li>
    <p><strong>Transparency matters for adoption</strong>
The traffic-light interface and editable feedback make the system easier for teachers to understand and trust.</p>
  </li>
  <li>
    <p><strong>AI works best as support, not replacement</strong>
Keeping the teacher in control ensures that the system augments professional judgment instead of obscuring it.</p>
  </li>
</ol>

<p>A key limitation is that the current validation is still preliminary: it was performed on a relatively small dataset from a single course context. Future work should test whether the same benefits hold for longer answers and other modalities such as <strong>code</strong>, <strong>mathematics</strong>, or other domain-specific responses.</p>

<p><strong>Future Directions</strong>:</p>

<ul>
  <li>Extending the cascade to longer and more complex responses</li>
  <li>Evaluating the approach across different subjects and answer modalities</li>
  <li>Studying how teachers use and adapt the feedback in real grading workflows</li>
  <li>Exploring richer semi-automatic interfaces that further reduce workload while preserving oversight</li>
</ul>

<hr />

<p><em>This blog post is based on research presented at the 3rd Workshop on Engineering Interactive Systems Embedding AI Technologies, 2025. For complete technical details, see the <a href="http://hdl.handle.net/1942/46328">full paper</a>.</em></p>]]></content><author><name>Jarne Thys</name><email>jarne.thys -at- uhasselt.be</email></author><category term="AI in Education" /><category term="Machine Learning" /><category term="Text Classification" /><summary type="html"><![CDATA[How cascaded transformers and a transparent Mixture-of-Agents interface improve educational grading while preserving teacher review and control.]]></summary></entry><entry><title type="html">INSIGHT: Bridging the Student-Teacher Gap in Times of Large Language Models</title><link href="https://jarnethys.com/ai%20in%20education/llms/educational%20technology/insight-bridging-student-teacher-gap/" rel="alternate" type="text/html" title="INSIGHT: Bridging the Student-Teacher Gap in Times of Large Language Models" /><published>2025-04-24T00:00:00+00:00</published><updated>2025-04-24T00:00:00+00:00</updated><id>https://jarnethys.com/ai%20in%20education/llms/educational%20technology/insight-bridging-student-teacher-gap</id><content type="html" xml:base="https://jarnethys.com/ai%20in%20education/llms/educational%20technology/insight-bridging-student-teacher-gap/"><![CDATA[<blockquote>
  <table>
    <tbody>
      <tr>
        <td><strong>Related Publication:</strong> <a href="/publication/24/04/2025-insight">INSIGHT: Bridging the Student-Teacher Gap in Times of Large Language Models</a></td>
        <td><a href="https://ceur-ws.org/Vol-4051/paper4.pdf">Full Paper</a></td>
      </tr>
    </tbody>
  </table>
</blockquote>

<p>This blog post describes research presented at the D-SAIL Workshop on Transformative Curriculum Design in Palermo, Italy (2025). Co-authored with S. Vanbrabant, D. Vanacken, and G. Rovelo Ruiz.</p>

<h2 id="what-we-did">What We Did</h2>

<p>This work focused on a problem many teaching staff now face: if students increasingly ask their questions to general-purpose LLMs instead of to instructors or teaching assistants, teachers lose an important source of feedback about where students are struggling. That makes it harder to adapt course materials, prepare targeted support, and maintain meaningful student-teacher interaction. INSIGHT was designed as a response to that problem.</p>

<p>The core goal was not to block AI use, but to create a <strong>privacy-aware, human-centered environment</strong> where students can use an LLM while their interactions still generate useful teaching insights. In other words, we wanted to explore how AI could help <em>bridge</em> the student-teacher gap instead of widening it.</p>

<p>This research builds directly on the ideas from my master’s thesis. In that earlier work, we explored how AI could support students and teaching staff with exercises through features such as LLM access, a dynamic FAQ, and teacher-side analytics. INSIGHT takes that direction further in a more focused proof of concept: it emphasizes modular deployment, explicit privacy choices, and the use of student questions as a signal for improving face-to-face support.</p>

<p>The paper identifies a clear tension in educational AI:</p>

<p><strong>Opportunities</strong>:</p>

<ul>
  <li>AI can help personalize learning and teaching support</li>
  <li>Students can get quick answers while solving exercises</li>
  <li>Teaching staff can reduce repetitive question load</li>
  <li>Interaction data can reveal patterns in student difficulties</li>
</ul>

<p><strong>Challenges</strong>:</p>

<ul>
  <li>Student-teacher interaction may degrade if too many questions move to AI</li>
  <li>Teaching staff may miss cues about knowledge gaps</li>
  <li>Privacy becomes a major concern when educational data is collected</li>
  <li>Students may become over-reliant on AI tools if they are used uncritically</li>
</ul>

<p>So the main research contribution was to design and present a modular prototype that gives students access to an LLM in a monitored, privacy-aware setting while equipping teaching staff with data-driven insights they can use to improve in-person teaching.</p>

<h2 id="how-we-did-it">How We Did It</h2>

<p>The project started with <strong>interviews and focus-group style discussions with teaching staff</strong> at Hasselt University. This happened in two phases: first through semi-structured interviews across different courses, and later by presenting an initial prototype during an internal university workshop focused on AI. These conversations shaped the design priorities of INSIGHT. The strongest recurring concern was that staff feared losing contact with students and, with that, losing insight into students’ understanding of the course material.</p>

<h3 id="system-design-insight">System Design: INSIGHT</h3>

<p>We developed INSIGHT as a <strong>modular proof-of-concept system</strong> centered around an <strong>INSIGHT Core</strong>. As shown in the architecture diagram on page 3, this core combines a keyword extraction component, a sentence similarity component, and a course database, and connects them to both a local LLM and a user interface that can be adapted to different course contexts.</p>

<h4 id="core-components">Core Components:</h4>

<ol>
  <li>
    <p><strong>AI-Assisted Exercise Support</strong></p>

    <ul>
      <li>Students can ask questions to an LLM while working on exercises</li>
      <li>The LLM is the main interaction method in the prototype</li>
      <li>In this implementation, we used <strong>Llama 3.2 3B</strong>, mainly because it offers relatively fast responses and is feasible to run on consumer hardware</li>
      <li>The modular setup makes it possible to swap in newer or more capable models later, depending on the needs of the course</li>
    </ul>
  </li>
</ol>

<p>A key design choice is that INSIGHT does <strong>not</strong> interfere with the LLM’s behavior through prompt engineering or artificial restrictions. The idea is that students should interact with it naturally, rather than being nudged into a constrained workflow. The system then analyzes those interactions afterward to generate useful signals for teaching staff.</p>

<ol>
  <li>
    <p><strong>Keyword Extraction</strong></p>

    <ul>
      <li>INSIGHT first needs to know which topics belong to each exercise</li>
      <li>To do that, it uses <strong>KeyBERT</strong> to extract an initial keyword list from the exercise text</li>
      <li>Teaching staff then review and refine those keywords manually in a mixed-initiative process</li>
      <li>These reviewed keywords become the course’s topic vocabulary</li>
    </ul>
  </li>
</ol>

<p>The example shown in the keyword selection interface on page 4 illustrates why human review matters: automatic extraction is helpful, but imperfect. The paper explicitly notes odd outputs such as “tree tree” in the extracted keywords for one exercise, which demonstrates why final control remains with the teaching staff.</p>

<ol>
  <li>
    <p><strong>Dynamic FAQ Generation</strong></p>

    <ul>
      <li>Student questions are embedded in the same vector space as the topic vocabulary</li>
      <li>The nearest keyword is used to assign a topic label to a question</li>
      <li>Similar questions are grouped using the <strong>all-MiniLM-L6-v2</strong> sentence embedding model and cosine similarity</li>
      <li>Recurrent questions can then be turned into FAQ items</li>
      <li>INSIGHT caches the LLM’s answer for grouped questions, and teaching staff can edit that answer before reuse</li>
    </ul>
  </li>
</ol>

<p>This makes the FAQ both dynamic and supervised. It grows from real student questions, but the teaching staff still verify and adjust the answers, which reduces variability and keeps the system human-in-the-loop.</p>

<ol>
  <li>
    <p><strong>Teaching Staff Insights Dashboard</strong></p>

    <ul>
      <li>INSIGHT collects and analyzes student questions to the LLM</li>
      <li>It also tracks FAQ usage and exercise difficulty ratings</li>
      <li>The teaching staff interface includes a dynamic FAQ and two visualizations</li>
      <li>
        <p>According to the interface shown on page 5, these graphs display:</p>

        <ul>
          <li>how often each FAQ item is viewed</li>
          <li>how frequently different course topics appear in student questions</li>
        </ul>
      </li>
    </ul>
  </li>
</ol>

<p>These visualizations give teaching staff a fast way to spot recurring confusion, identify weak points in course materials, and prepare more targeted support during office hours or interactive sessions.</p>

<h3 id="modular-design">Modular Design</h3>

<p>INSIGHT is designed to be:</p>

<ul>
  <li><strong>Flexible</strong>: The core system is separate from the user interface and LLM, so it can be deployed in different higher education courses without changing the core logic. The proof of concept uses data from UHasselt’s <em>Algorithms and Data Structures</em> course, but the architecture is intentionally broader than that one course.</li>
  <li><strong>Privacy-aware</strong>: Students can choose an <strong>anonymous mode</strong>, allowing staff to see the data without attached names. In addition, the system uses <strong>Ollama</strong> to run the LLM locally, so queries can remain private and are not shared with third parties.</li>
  <li><strong>Extensible</strong>: Because the components are modular, new models, interfaces, and analytics features can be added later without redesigning the full system.</li>
</ul>

<p>This privacy design is especially important because the system deliberately collects interaction data. INSIGHT’s approach is to make that collection transparent and consent-based, rather than hidden or mandatory.</p>

<h3 id="implementation">Implementation</h3>

<p>Technically, the system consists of:</p>

<ul>
  <li>a central <strong>INSIGHT Core</strong></li>
  <li>a <strong>database</strong> storing course information and interaction data</li>
  <li>a <strong>keyword extraction module</strong></li>
  <li>a <strong>sentence similarity module</strong></li>
  <li>a <strong>local LLM</strong></li>
  <li>separate <strong>teacher</strong> and <strong>student</strong> user interfaces</li>
</ul>

<p>The student interface, shown on page 6, combines three functions:</p>

<ul>
  <li>a chat window for asking the LLM questions</li>
  <li>access to the FAQ</li>
  <li>an exercise difficulty rating mechanism that gives explicit feedback to the teaching staff</li>
</ul>

<p>The teacher interface, shown on page 5, allows staff to:</p>

<ul>
  <li>review and edit FAQ items</li>
  <li>add FAQ entries manually</li>
  <li>inspect topic frequencies</li>
  <li>inspect FAQ view counts</li>
</ul>

<h2 id="results">Results</h2>

<p>As a proof of concept, INSIGHT demonstrates that it is possible to combine LLM access, keyword analysis, dynamic FAQs, and teacher-facing analytics into a single educational support system that strengthens visibility into student difficulties instead of hiding them.</p>

<h3 id="system-capabilities">System Capabilities</h3>

<p>INSIGHT successfully demonstrated:</p>

<ol>
  <li>
    <p><strong>Better Student-Teacher Connection</strong>
The system is explicitly designed so that student questions to AI still become useful signals for the teaching staff. Rather than replacing face-to-face support, it helps staff understand where students need help before or during in-person interactions.</p>
  </li>
  <li>
    <p><strong>Data-Driven Personalization</strong>
By analyzing question topics, FAQ usage, and difficulty ratings, teaching staff can:</p>

    <ul>
      <li>identify which topics cause recurring confusion</li>
      <li>adapt course materials over time</li>
      <li>prepare more targeted explanations for students or groups</li>
      <li>provide more personalized face-to-face support</li>
    </ul>
  </li>
  <li>
    <p><strong>Dynamic FAQ with Human Oversight</strong>
Frequently recurring questions can be grouped, answered once, and then reviewed by staff. This reduces repetitive question load while still allowing teachers to verify accuracy and refine explanations.</p>
  </li>
  <li>
    <p><strong>Privacy-Conscious Monitoring</strong>
The system shows that monitoring does not have to mean intrusive surveillance: students can opt into identified sharing or choose anonymous use, and local deployment reduces third-party data exposure.</p>
  </li>
</ol>

<h3 id="insights-from-deployment">Insights from Deployment</h3>

<p>This paper presents INSIGHT primarily as a <strong>proof of concept</strong>, not a full classroom validation study. The demonstrated value lies in the kinds of insights the system can surface for teaching staff, especially in interactive course settings.</p>

<p>Based on the prototype, staff can derive insights such as:</p>

<ul>
  <li>which exercise topics are producing the most questions,</li>
  <li>which questions recur often enough to justify an FAQ item,</li>
  <li>which FAQ items students revisit frequently,</li>
  <li>and which exercises students rate as more difficult.</li>
</ul>

<p>Those signals can help answer practical teaching questions like:</p>

<ul>
  <li><em>Which concept needs clearer explanation in the next session?</em></li>
  <li><em>Which exercise caused unexpected confusion?</em></li>
  <li><em>Which FAQ answers should be strengthened or expanded?</em></li>
</ul>

<p>The paper also discusses important limitations. Students may avoid monitored systems and ask questions to outside LLMs instead, which could bias the teacher’s view of student understanding. That means the usefulness of INSIGHT depends not only on the tool itself, but also on whether students perceive value in using it and whether teaching staff actively act on the insights it provides.</p>

<h2 id="conclusion">Conclusion</h2>

<p>INSIGHT shows that AI in education does not have to reduce human contact by default. If designed carefully, it can create a <strong>shared space</strong> where students get quick AI support and teaching staff still retain visibility into learning difficulties, common misconceptions, and emerging needs.</p>

<p><strong>Key Contributions</strong>:</p>

<ol>
  <li>A <strong>modular architecture</strong> for integrating AI support into higher education courses</li>
  <li>A <strong>mixed-initiative keyword analysis pipeline</strong> to map student questions to course topics</li>
  <li>A <strong>dynamic FAQ system</strong> built from real student questions and supervised by teaching staff</li>
  <li>A <strong>teacher-facing analytics layer</strong> that supports more personalized face-to-face teaching</li>
  <li>A <strong>privacy-aware deployment model</strong> with anonymous mode and local LLM execution</li>
</ol>

<p>What stands out most in this work is the design philosophy: INSIGHT does not try to stop students from using LLMs, nor does it assume AI should replace teachers. Instead, it accepts that these tools are already part of students’ reality and tries to make their use safer, more transparent, and more useful for actual teaching practice.</p>

<p>The paper also makes clear that technical design alone is not enough. Responsible AI use in education depends on dialogue, trust, and shared expectations between students and staff. INSIGHT supports that conversation, but it cannot replace it.</p>

<p><strong>Future Work</strong>:</p>

<p>The next steps described in the paper are:</p>

<ul>
  <li>conducting a <strong>cognitive walkthrough</strong> with new teaching staff interested in piloting the system,</li>
  <li>deploying INSIGHT in <strong>real courses at Hasselt University</strong>,</li>
  <li>empirically validating whether it improves support and learning,</li>
  <li>exploring <strong>knowledge tracing</strong> for more advanced personalization,</li>
  <li>and using the collected data to support <strong>adaptive learning</strong> and more inclusive learning experiences over time.</li>
</ul>

<p>If we were to extend this line of work, the most important next step would be a real classroom study: not just to test whether the system works technically, but to understand whether students actually use it, whether teachers trust it, and whether it genuinely improves student-teacher interaction in practice.</p>

<hr />

<p><em>This blog post is based on research presented at the D-SAIL Workshop, 2025. For complete details, see the <a href="https://ceur-ws.org/Vol-4051/paper4.pdf">full paper</a>.</em></p>]]></content><author><name>Jarne Thys</name><email>jarne.thys -at- uhasselt.be</email></author><category term="AI in Education" /><category term="LLMs" /><category term="Educational Technology" /><summary type="html"><![CDATA[A research account of designing INSIGHT to turn student LLM questions into teaching insights while preserving privacy, agency, and human support.]]></summary></entry><entry><title type="html">Developing an AI-Based Prototype to Support Students and Teachers with Exercises</title><link href="https://jarnethys.com/ai%20in%20education/master's%20thesis/ai-education-master-thesis/" rel="alternate" type="text/html" title="Developing an AI-Based Prototype to Support Students and Teachers with Exercises" /><published>2024-09-13T00:00:00+00:00</published><updated>2024-09-13T00:00:00+00:00</updated><id>https://jarnethys.com/ai%20in%20education/master&apos;s%20thesis/ai-education-master-thesis</id><content type="html" xml:base="https://jarnethys.com/ai%20in%20education/master&apos;s%20thesis/ai-education-master-thesis/"><![CDATA[<blockquote>
  <table>
    <tbody>
      <tr>
        <td><strong>Related Publication:</strong> <a href="/publication/2024-09-13-masters-thesis">Ontwikkeling van een AI-gebaseerd prototype dat studenten en docenten ondersteunt met oefeningen</a></td>
        <td><a href="http://hdl.handle.net/1942/43873">Full Thesis</a></td>
      </tr>
    </tbody>
  </table>
</blockquote>

<p>This blog post describes my master’s thesis work on using AI in education, supervised by Davy Vanacken and Gustavo Rovelo Ruiz at UHasselt.</p>

<h2 id="what-i-did">What I Did</h2>

<p>My thesis explored how AI—especially Large Language Models (LLMs)—can be integrated into education in a way that is useful, transparent, and responsible.</p>

<p>With the emergence of LLMs, students now have the opportunity to ask AI questions. However, teachers have concerns about the impact on teaching quality—particularly regarding the accuracy of AI-generated answers, the consistency of those answers, privacy, and the potential reduction in student-teacher interaction.</p>

<p>The goal of this thesis was to:</p>

<ul>
  <li>Explore applications and challenges of AI (especially LLMs) in education</li>
  <li>Create a prototype demonstrating how these applications can be implemented</li>
  <li>Ensure that AI enhances rather than replaces student-teacher interaction</li>
</ul>

<p>More specifically, the project was structured around three objectives:</p>

<ol>
  <li><strong>Identify useful AI tools and their challenges in education</strong> through a literature review and exploratory experiments with LLM-generated answers discussed with teachers.</li>
  <li><strong>Develop a student-facing prototype</strong> that gives responsible access to an LLM, offers a dynamic FAQ, and recommends exercises based on earlier difficulties.</li>
  <li><strong>Develop a teacher-facing prototype</strong> that helps teachers manage the LLM, review frequently asked questions, inspect learning difficulties, and use data visualisations to improve course materials.</li>
</ol>

<p>The broader research question was not whether AI should replace teachers, but how it can be integrated in a way that supports both students and teaching teams while preserving human oversight.</p>

<h2 id="how-i-did-it">How I Did It</h2>

<p>The research followed a structured process inspired by the ADDIE model: analysis, design, development, implementation, and evaluation.</p>

<ol>
  <li>
    <p><strong>Literature Review</strong>
I conducted a broad review of existing AI applications in education, including machine learning, personalised learning systems, chatbots, expert systems, intelligent tutors, and virtual learning environments. I also focused on the main challenges of LLM use in education: unreliable answers, inconsistent outputs, issues around assessment, privacy, intellectual property, and the risk of weakening teacher-student interaction.</p>
  </li>
  <li>
    <p><strong>Experiments with LLMs</strong>
Before designing the prototype, I explored how LLMs behaved when solving exercises and discussed those outputs with teachers. This helped clarify where LLMs were useful, where they failed, and what kinds of controls teachers would need to trust them in an educational setting. This phase also informed the role of prompt engineering in steering LLM responses toward course-specific expectations.</p>
  </li>
  <li>
    <p><strong>Concept Development</strong>
Based on the literature and experiments, I defined four core design goals for the prototype:</p>

    <ul>
      <li>Give students responsible access to an LLM</li>
      <li>Build a dynamic FAQ that evolves from student questions</li>
      <li>Recommend exercises based on earlier difficulties</li>
      <li>Make the whole application transparent about which parts use AI and why</li>
    </ul>
  </li>
  <li>
    <p><strong>Prototype Implementation</strong>
I built two applications using <strong>Windows Presentation Foundation (WPF)</strong> with <strong>C#</strong>:</p>

    <ul>
      <li>one for <strong>teachers</strong></li>
      <li>one for <strong>students</strong></li>
    </ul>

    <p>Both applications communicate with a central <strong>Python Flask server</strong>, which acts as the hub between the interfaces and the AI/ML components.</p>

    <p>The system combines several components:</p>

    <ul>
      <li>An <strong>open-source LLM</strong> for student questions</li>
      <li><strong>KeyBERT</strong> for keyword extraction from exercises and questions</li>
      <li><strong>Sentence Transformers</strong> for grouping similar questions using sentence similarity</li>
      <li>A <strong>collaborative filtering recommender</strong> for exercise suggestions</li>
      <li>A <strong>database</strong> for courses, questions, ratings, and related interaction data</li>
    </ul>
  </li>
  <li>
    <p><strong>Design Decisions</strong>
A key design choice was to rely on <strong>locally manageable, open-source models</strong> rather than a fully external commercial AI provider. In the prototype, the default model was <strong>GPT4All Falcon</strong>, chosen because it can run on relatively modest hardware and gives teachers more control. Teachers can also configure course-specific model URLs through Hugging Face-compatible APIs and provide custom instructions such as an opening prompt and instructions that are appended to every student message.</p>
  </li>
</ol>

<h2 id="results">Results</h2>

<p>The prototype demonstrated how multiple AI-assisted features can be combined into a practical educational workflow.</p>

<h3 id="teacher-side-results">Teacher-side results</h3>

<p>The teacher application gives control over:</p>

<ul>
  <li><strong>Which LLM students use</strong></li>
  <li><strong>How the LLM behaves</strong>, through course-specific instructions</li>
  <li><strong>Which keywords are associated with each exercise</strong></li>
  <li><strong>Which FAQ items are added</strong></li>
  <li><strong>Which trends and difficulties appear in student data</strong></li>
</ul>

<p>Teachers can:</p>

<ul>
  <li>Add a course and configure the LLM</li>
  <li>Input exercises and curate automatically extracted keywords</li>
  <li>Review grouped student questions</li>
  <li>Receive <strong>suggestions for new FAQ items</strong></li>
  <li>Inspect <strong>exercise completion rates</strong></li>
  <li>View <strong>perceived exercise difficulty</strong></li>
  <li>Use visualisations to better understand where students struggle</li>
</ul>

<p>This makes the AI layer adjustable rather than opaque.</p>

<h3 id="student-side-results">Student-side results</h3>

<p>The student application focuses on guided support:</p>

<ul>
  <li>Students can chat with the LLM in a familiar interface</li>
  <li>They can select a course and optionally a specific exercise, which links questions to the right context</li>
  <li>They can browse a <strong>dynamic FAQ</strong> with previously answered questions and teacher-added hints</li>
  <li>They can receive <strong>exercise recommendations</strong> based on earlier ratings and difficulty patterns</li>
</ul>

<p>An important addition is the option to use the system in an <strong>anonymous mode</strong>. This allows students to contribute data about difficult topics and common questions without attaching that data to their identity, while clearly warning them that anonymous use limits the teacher’s ability to give individual feedback.</p>

<h3 id="core-technical-outcomes">Core technical outcomes</h3>

<p>The prototype successfully demonstrated several innovations:</p>

<ul>
  <li><strong>Teacher-Customizable LLMs</strong>: Teachers can configure AI responses to align with their teaching style and course requirements</li>
  <li><strong>Exercise Recommender</strong>: A collaborative filtering system predicts which exercises may help a student practise weak areas</li>
  <li><strong>Dynamic FAQ</strong>: Student questions are grouped through sentence similarity, and commonly recurring questions can be proposed to teachers as FAQ additions</li>
  <li><strong>Data Collection for Insights</strong>: Questions, ratings, completion data, and detected keywords help teachers understand learning bottlenecks</li>
</ul>

<p><strong>Key Insight</strong>: Rather than diminishing interaction between teachers and students, the collected data can help teachers maintain and even improve that interaction by making student difficulties more visible and by supporting more targeted feedback.</p>

<p>At the same time, the thesis also identified important limitations:</p>

<ul>
  <li>LLM reliability is still a real issue</li>
  <li>The chosen LLM was efficient, but limited (for example, it only responded in English)</li>
  <li>The recommender suffers from the classic <strong>cold start problem</strong>, since recommendations are weaker when little prior data is available</li>
  <li>Some desirable features, such as stronger emotional intelligence or deeper reflection, would require far more computational resources or further model adaptation</li>
</ul>

<p>Because this thesis focused on concept validation and implementation, it did <strong>not</strong> yet include a full user study with classroom participants. That evaluation is the most important next step.</p>

<h2 id="conclusion">Conclusion</h2>

<p>This work shows that AI in education does not have to mean replacing teachers. Instead, it can be used to create a more supportive and data-informed learning environment—provided that the system is transparent, locally manageable, and designed around teacher control.</p>

<p>The thesis started with a broad analysis of AI opportunities and risks in education, then translated those findings into a working prototype for both students and teachers. The result is a concrete demonstration of how LLMs, keyword extraction, sentence similarity, and recommender systems can be combined in an educational setting while addressing important concerns such as transparency, privacy, and meaningful human oversight.</p>

<p>Several lessons stood out during the project:</p>

<ul>
  <li><strong>Control matters</strong>: teachers need ways to shape and constrain AI behaviour</li>
  <li><strong>AI works best as a complement</strong>: not as a replacement for teacher-student interaction</li>
  <li><strong>Transparency increases trust</strong>: users should understand which parts are AI-driven</li>
  <li><strong>Local deployment helps privacy</strong>: avoiding unnecessary external data sharing is a major advantage</li>
</ul>

<p><strong>Future Work</strong>: The most important next step is a real user study comparing a control group and a test group over a longer period. Such a study should evaluate learning outcomes, retention, usability, transparency, student autonomy, and whether teachers still feel in control of the learning process. There is also room to improve the recommender, add more prompt-engineering support for teachers, and explore additional AI applications such as AI-assisted course design or tutoring.</p>

<p>If I were to extend this work, I would focus first on evaluating the prototype with real users, then on refining the recommendation logic and making the LLM support more accessible for teachers with little AI experience.</p>

<hr />

<p><em>This blog post is based on my master’s thesis completed at UHasselt in 2024. For the complete technical details, please refer to the <a href="http://hdl.handle.net/1942/43873">full thesis</a>.</em></p>]]></content><author><name>Jarne Thys</name><email>jarne.thys -at- uhasselt.be</email></author><category term="AI in Education" /><category term="Master&apos;s Thesis" /><summary type="html"><![CDATA[How a transparent human-centered AI for education prototype supports student exercises, teacher insight, and continued student-teacher interaction.]]></summary></entry></feed>