Interpretive Learning Analytics: The Dashboard Knows Less

From Activity Traces To AI-Driven Decisions

Imagine a learner completing a short online assessment at the end of a training module. Ten questions, eight correct answers. Nothing unusual so far. But those eight answers may not remain just an assessment result. They can raise a proficiency score, alter a learning path, remove the learner from additional support, and eventually contribute to a judgment about readiness for a role.

Eight correct answers can travel surprisingly far.

They can raise a proficiency score, alter a learning path, remove a learner from support, and eventually contribute to a judgment about readiness for a role. Every step may appear reasonable. The underlying data may even be perfectly accurate. Yet the final conclusion can claim far more than the evidence supports.

What the system has actually observed is narrow: 8 answers submitted correctly, perhaps after 15 minutes into a module. What it may eventually represent is much larger: that the learner understands the subject, possesses a skill at a particular level, can perform independently, and is ready to move on. Between those two statements lies one of the least visible parts of learning analytics.

That interpretive step is not external to learning analytics; it is part of the field itself. SoLAR’s 2025 definition explicitly describes learning analytics as the collection, analysis, interpretation, and communication of data about learners and their learning. The important word here is interpretation. Data do not arrive with their educational meaning already attached.

AI is making that hidden passage faster and more consequential. A tentative signal can now update a profile, generate a pathway, trigger an enrollment, remove a support intervention, or travel into another system before anyone asks what the original evidence actually justified. The central question is therefore not how much learning data an organization can collect. It is how much meaning—and how much consequence—it is entitled to attach to them.

The Data Can Be Right, And The Conclusion Still Be Weak

Discussions about learning data often begin with data quality. Are records complete? Are timestamps reliable? Are activities correctly attributed? Those questions matter, but they are only the first layer.

Imagine that the data are flawless. A learner really did answer eight questions correctly. The platform really did record 15 minutes between 2 events. The module really was completed according to its rules. Nothing is inaccurate.

An 80% score is strong evidence that eight of ten answers were correct under those conditions. Depending on the assessment design, it may also be useful evidence for offering more advanced material. It does not automatically demonstrate durable understanding, independent performance, or readiness for a consequential decision.

Time-on-task makes the problem especially visible. If one event is recorded at 10:02 and the next at 10:17, the system knows that 2 events occurred 15 minutes apart. It does not know what happened throughout those 15 minutes. A study by Kovanović and colleagues showed that different methods for estimating time-on-task can materially change learning analytics findings [1].

The issue is therefore not simply reliability—whether a measurement is recorded or applied consistently—but validity: whether the evidence actually supports the interpretation and use attached to it. Philip Winne’s analysis of trace-based learning analytics makes the same point more formally: theory is unavoidable when deciding what should be observed and what inferences and actions those observations can justify.

Precision of presentation is not strength of evidence. A completion rate of 100%, a proficiency score of 82%, or a label such as „advanced“ can look like facts. But the metric becomes misleading when it is asked to answer a question it was never designed to answer.

The Hidden Chain Behind The Dashboard

Between an activity performed by a learner and an action taken by a system lies a chain that dashboards rarely make fully visible:

activity trace → metric → inference → profile → decision/action

A trace records something observable: a resource was opened, an answer submitted, a video paused, an activity completed. Technical standards make this distinction surprisingly clear. The current ISO/IEC/IEEE xAPI standard is designed to communicate experiential activity data to a Learning Record Store, while 1EdTech Caliper [2] defines common ways to capture and describe learning activity and product-usage data. These infrastructures can record that something happened. They do not, by themselves, establish what that event means for learning.

A metric organizes traces: 80% correct, 12 minutes, 3 attempts, 5 resources viewed. An inference then attaches meaning to that metric: the learner understands, is struggling, is engaged, has mastered a skill, or is ready for a more difficult task. A profile stabilizes that inference as an attribute associated with a person. Finally, a decision does something with the representation: recommend content, change a pathway, send an alert, remove support, or update another system.

The first stage describes an event. The later stages increasingly describe a person.

Every transition contains a design choice. Which events count? How are they aggregated? What threshold separates „basic“ from „advanced“? What alternative explanations fit the same pattern? How long does an inferred attribute remain valid? Who is allowed to act on it?

The system has not discovered competence in the same way it recorded a submitted answer. It has applied a model according to which a particular pattern is treated as evidence of competence. That model may be reasonable. It is still a model.

When A Hypothesis Becomes A Profile

Profiles are persuasive because they are built from nouns: skill, interest, gap, strength, readiness, proficiency. Nouns look stable. The verbs that produced them are harder to see: observed, aggregated, weighted, compared, predicted, inferred.

A profile might display „Data Analysis — Advanced.“ What it rarely displays with equal prominence is „inferred from two assessment scores, three completed activities, and a weighting rule applied six months ago.“ Yet provenance changes how confidently the information should be used.

This is not merely a philosophical concern about language. A 2023 review of Learning Analytics and Knowledge conference papers and Journal of Learning Analytics articles found that 71.1% of the empirical articles examined did not include any measure of learning outcomes. The result does not show that learning analytics cannot measure learning. It does show why learner data, activity data, and evidence of learning should not be treated as interchangeable categories.

Once an inference enters a profile, it can also begin shaping the evidence that comes next. A learner classified as advanced receives different opportunities. A learner marked as needing support receives another pathway. Future activity then occurs inside an environment partly structured by earlier classifications, producing new data that feed the profile again. A hypothesis can therefore become self-reinforcing if its origin, uncertainty, and reversibility disappear from view.

AI Does Not Create The Leap. It Removes The Pause.

Rules-based systems have long turned metrics into actions: score below 60%, assign remediation; complete a module, unlock the next one. AI does not invent the move from inference to consequence.

What changes is the speed, scale, and opacity with which several steps can be combined. A system can analyze a profile, identify an apparent gap, generate an assessment, evaluate the response, update the profile, select the next resource, and trigger another workflow. The distance between „the model thinks“ and „the system acts“ can become very short.

That matters because the quality of an inference cannot be separated from the authority granted to it. NIST’s AI Risk Management Framework treats validity, reliability, accountability, transparency, explainability, and interpretability as context-dependent characteristics of trustworthy AI [3]. In learning systems, context includes the consequence attached to the output.

Jisc’s Code of Practice for Learning Analytics reaches a compatible conclusion from the learning-data side: institutions should make clear the purposes, data sources, metrics, processes, boundaries of use, and interventions associated with analytics. In other words, the chain should be legible before it becomes consequential.

An AI recommendation leaves room for a pause: the learner or manager can accept it, reject it, or ask for another view. An automated action can remove that pause. Once a system is authorized to act on its own conclusions, „How good is this inference?“ must be asked together with „What is this inference allowed to do?“

Match The Evidence To The Consequence

The same evidence can be adequate for one purpose and inadequate for another. Eight correct answers may be enough to suggest an optional advanced resource, particularly if the learner can ignore the recommendation or return to an earlier activity. They may be nowhere near enough to certify competence, withdraw necessary support, alter an official professional profile, or contribute to readiness for a role.

A useful design principle follows: the strength and diversity of the evidence should rise with the weight of the consequence.

This does not mean collecting more data indiscriminately. More traces can simply produce a larger collection of the same weak evidence. What matters is whether the evidence is appropriate to the conclusion. If the question is whether an activity was completed, completion data may be sufficient. If the question is whether a learner can perform independently, the evidence must make independent performance observable. If a decision affects access or opportunity, the system needs stronger grounds, meaningful human review, and a clear route for correction. Before turning a learning metric into action, L&D teams should be able to answer five questions:

  1. What event was actually observed, and what remained outside the system’s view?
  2. Why was this trace transformed into this particular metric?
  3. Which alternative explanations could produce the same pattern?
  4. How strong is the inference for this specific purpose and context?
  5. What consequence will follow, and who can review, correct, or reverse it?

These questions do not prevent automation. They make automation proportionate.

Keep Evidence From Becoming A Verdict

Learning analytics can reveal patterns, direct attention, improve resource design, and support better decisions. AI can connect information across systems and respond faster to emerging needs. The answer is not to reject either capability. The answer is to preserve the distinction between observation, inference, and decision.

A mature analytics environment should make important conclusions traceable to their evidence, keep uncertainty visible when uncertainty is real, and make profiles open to revision. It should also recognize that a metric suitable for guidance may not be suitable for judgment.

A dashboard should help us decide where to look. It should not quietly decide what counts as knowing. Eight correct answers can remain eight correct answers: useful, interpretable evidence that may justify a next question, a recommendation, or further assessment. They do not have to become a verdict on the learner.

The promise of AI-driven learning analytics is not that systems will finally know learners completely. It is that they can help us notice more, connect more evidence, and make better decisions—provided we can still see the difference between what was observed, what was inferred, and what we decided to do about it.

References:

[1] Does Time-on-task Estimation Matter? Implications on Validity of Learning Analytics Findings

[2] Caliper Analytics

[3] AI Risk Management Framework

Sources:

Schreibe einen Kommentar