1. Home
  2. Domain
The research

Why we built AURIVA

Everything behind the app in seven friendly chapters: what exists today, the problem, the gap, our goals, how we worked, what we used and what we found. Open the "for researchers" notes for the technical detail, or read each full thesis.

Literature survey

We started by studying the apps that autistic learners and their teachers already use: Read Along, Proloquo, Avaz AAC, Otsimo, MITA, AutiSpark, Writing Wizard and LetterSchool. They are good at pictures, sounds and communication. We compared what each one offers against what a classroom really needs.

What the app offersRead AlongProloquoAvaz AACOtsimoMITAAutiSparkWriting WizardLetter SchoolAURIVA
Designed for autistic learnersβœ˜βœ”βœ”βœ”βœ”βœ”βœ˜βœ˜βœ”
Pictures + sound + touchβœ”βœ”βœ”βœ”βœ”βœ”βœ”βœ”βœ”
Structured pronunciation learningβœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ”
Records speech and scores itβœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ”
Adapts to each child's progressLimited✘✘LimitedLimitedLimitedβœ˜βœ˜βœ”
Teacher progress dashboard✘LimitedLimitedLimited✘✘Limitedβœ˜βœ”
Handwriting motor analysisβœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ”
Conversation practiceβœ˜βœ”βœ”βœ˜βœ˜βœ˜βœ˜βœ˜βœ”
Matches the Sri Lankan curriculumβœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ˜βœ”

Summary of the literature review in our project proposal (March 2026). "Limited" means the feature is only partly available. Each thesis then reviewed its own area in depth, as below.

What each module's literature review found

Concept learning

Hands-on audits of two popular apps (Otsimo and Speech Blub) found no tiered mastery, no graduated "try again" help and no model of how concepts relate. Many AI studies report very high accuracy, but on old datasets or with very few, non-autistic testers.

Handwriting

Handwriting difficulties in some autistic children link to fine motor control. How a letter is written (path, timing, pauses) is studied mostly for screening, not used in everyday practice, and tracing apps judge only the final shape.

Pronunciation

Computer-assisted pronunciation tools are mature, but almost all assume adult or typically developing learners. Child speech is harder for machines, and autism-focused tools rarely teach pronunciation at all.

Dialogue

Vocabulary tools test whether a child can name a word, not when to use it. Common "forgiving" speech matching also rewards mistakes such as "tank you", and echolalia (repeating without understanding) can be mistaken for learning.

Research problem

Many autistic children in Sri Lanka find it hard to understand language and learn new concepts. English is a core school subject, yet they get almost no specialised support for it. Ordinary teaching lacks the clear visual structure and step-by-step help that they need, and teachers who look after many different learners have no easy way to see who needs what.

1 in 127children diagnosed with autism worldwide
1 in 93children affected in Sri Lanka
9,000identified children with autism in Sri Lanka
<15%learn in genuinely inclusive classrooms

Of these children, about 23.5% are excluded from mainstream education and 8.6% attend schools specialised for disabled children.

Who feels this problem?

Autistic children

Harder to learn English, slower progress, and less confidence and engagement in class.

Special education teachers

Looking after many different learners at once, with no easy way to personalise lessons or follow each child's progress.

Principals & schools

No outcome data to prove inclusive education is working, and hard to justify resources without it.

Research gap

Existing tools treat words, writing, speaking and conversation as separate activities, and rarely give teachers a clear view across them. Our four theses each found a specific gap that no current tool fills for autistic learners in Sri Lanka:

Curriculum-matched lessons

Tools are not mapped to the Sri Lankan English syllabus, so teachers must fit global content to local classes.

How concepts relate

When a child mixes up "apple" and "cherry", that mistake has a pattern. No tool uses it, so each new word starts from nothing.

Writing process, not just the final shape

Apps check whether the traced letter looks right, not how the hand moved to make it, which is what teachers need to see.

Sound-by-sound pronunciation help

Pronunciation tools are built for adults. None are designed for autistic children or pinpoint which sound is hard.

Using words, not only naming them

Words are taught in isolation, speech checks are not fair to Sinhala-speaking children, and echoing is mistaken for learning.

Explainable evidence for teachers

Teachers need to see why a child was moved on or held back. A hidden score from a black box is not enough.

Research objectives

Main objective: to develop and evaluate a teacher-assisted, adaptive English learning platform that brings concept learning, handwriting, pronunciation and dialogue together for autistic children in Sri Lanka, keeping the teacher in charge of every learning decision.

Each team member owns one module with its own aim:

Concept learning

Design, deploy and evaluate a module that teaches curriculum vocabulary in three tiers with a mastery rule, uses a graph of concepts to understand confusions, and tests whether graph attention beats simpler rules.

Handwriting

Capture handwriting on a tablet, turn it into easy-to-read motor measures, run a transparent three-attempt practice flow for letters and words, and give teachers clear evidence.

Pronunciation

Help Grade 3 learners practise English sounds with pictures, audio and mouth-shape guidance, score speech sound by sound, choose the next word from the weakest sound, and keep teachers in the review loop.

Dialogue

Teach everyday words and short sentences, assess speech fairly for Sinhala speakers, detect echoing from response time, and predict each child's learning pace with an explainable model.

Methodology

We work in four stages, always with teachers and the school in the loop. Each researcher then chose a study design that fits their own question.

  1. ListenField study, literature review and requirement gathering
  2. DesignSystem architecture and child-friendly screens
  3. BuildFour modules, then joined into one platform
  4. Test & improveAutomated tests, classroom studies and panel feedback

Four modules, four studies

What it does

Each word is taught in three tiers: Find it (see the picture, hear the word, pick the match), Name it (choose the written word) and Watch and colour (a short video and a free colouring activity). A child must get two out of three right in a block to pass a tier. If they struggle, the help changes: first a cartoon version, then a short animation, then the full introduction again. After repeated difficulty the app hands over to the teacher instead of looping.

How it was studied

  • Design: mixed methods, comparing each child with their own earlier attempts (no control group, for ethical reasons).
  • Setting: 21 registered children at a special education school in Gampaha, over 27 active days (2 May to 9 October 2026).
  • Content: 96 concepts in 9 categories drawn from the National Institute of Education Grade 1 and 2 English guides.
  • Measures: tier mastery, learning gain, confusions between concepts, engagement, and teacher feedback.
For researchers: RQ1: does tier-gated learning produce measurable gain? RQ2: do visual or phonetic similarity predict real confusions? RQ3: does a Graph Attention Network over the Neo4j knowledge base beat the deployed heuristic? Evaluated with leave-one-student-out cross-validation against five classical baselines, run on versioned data snapshots so every number is reproducible. Mann-Whitney tests with effect sizes were used for the similarity analysis.
See the findings

What it does

Children write with a stylus on a tablet. AURIVA records the writing as a time-ordered path and compares it with a reference letter, looking at smoothness, pauses, direction control and speed. Each letter is practised in three attempts with less and less help, and mastery is decided from the third, unaided attempt (a score of 70 or more). There are at most three cycles per letter each day, then targeted help for that exact letter, a teacher-approved worksheet for extra practice at home, and finally words. Teachers get a report with the exact evidence and a "Writing Check".

How it was studied

  • Design: design-oriented mixed methods: the component was built step by step, then tested with data from controlled sessions on a stylus tablet at a partner school, and with teachers.
  • Data: 7 learners drew six pre-writing shapes (42 trajectories) and wrote 20 reference letters, three attempts each (420 letter records).
  • Checks: the research code was tested to reproduce the app's own numbers exactly, plus an automated test suite.
  • Machine learning: used only to describe patterns for researchers. It never decides what a child practises next.
For researchers: direction-invariant Dynamic Time Warping against a shared letter template; interpretable motor features combine into a weighted Motor Score; rule-based progression. Exploratory unsupervised analysis (K-Means, Ward clustering, Gaussian mixtures) is kept separate from operational decisions, with clusters treated as descriptive and not diagnostic. Stability was tested with 100 initialisations and leave-one-participant-out agreement (ARI).
See the findings

What it does

The teacher picks the alphabet or a word category (animals, fruits, classroom objects or everyday actions). The child sees a picture, hears a clear UK-English reference and sees mouth-shape guidance for each sound, then plays activities such as listen-and-choose, Tap the Sounds (no speaking needed) and spoken-word practice. The child sees only encouragement and never a number. The teacher sees the score, the weakest sound, a recording to replay and a queue of uncertain results to review. The next word is chosen from the weakest sound, the word's difficulty and repeated problems. In all, 78 items are covered: 26 letters and 52 words.

How it was studied

  • Design: design-and-development research, tested against teachers' own judgement.
  • Agreement test: 249 real scored attempts were recorded; teachers reviewed 32 of them (chosen where the app was least certain) and their scores were compared with the app's.
  • Child-speech adaptation: 167 child utterances were separated automatically from teacher prompts in 9 classroom sessions; teacher labelling for training is still in progress.
For researchers: a three-layer scoring cascade: (1) an accept-only MFCC and DTW comparison, (2) phoneme-level Goodness of Pronunciation with wav2vec2, backed by Whisper word checking, and (3) recalibration from teacher reviews that is applied only if validation shows it helps. MFCC-DTW alone proved unreliable for children's voices, which is why the phoneme layer was added. Adaptation is evaluated with student-held-out cross-validation.
See the findings

What it does

Level 1 teaches 30 everyday words in three groups: Greetings, Magic Words ("thank you", "excuse me") and Abilities ("clap", "eat", "sleep"). Each word has three phases: see and hear it, say it, and choose the situation where it fits. Level 2 builds short personalised sentences for the Grade 3 speaking topics: introducing yourself, describing a friend and describing a pet. A no-speech pathway runs through both levels so every child can take part.

Three research ideas

  • Fair speech scoring: speech is checked sound by sound, and typical Sinhala-to-English mix-ups get partial credit instead of a fail.
  • Echo detection: if a child answers almost instantly after hearing the word, it is probably an echo. It is not counted as learning, and the child is never penalised.
  • Explainable pace prediction: a two-step model predicts whether a child is moving at a fast, typical or struggling pace and tells the teacher why.

A design change we are proud of. The proposal included a third level where an AI generated whole conversations. Test rounds showed it was too hard and too unpredictable for the children, so we removed it and put the effort into making basic words and sentences comfortable first.

How it was studied

  • Design: Design Science Research with repeated test rounds and 114 passing unit tests.
  • Field study: supervised, with 10 children and about six teachers at three schools in the Colombo district, running since 7 October 2026.
For researchers: RQ1 compares phoneme-aware scoring with a human rater; RQ2 compares a Random Forest on first-session features against a literature-grounded formula using leave-one-child-out validation; RQ3 measures how accurately response latency identifies immediate echolalia; RQ4 looks at teacher usability. The learned model sits behind an interpretable formula, a confidence gate and a switch, so it can never change a child's recorded progress.
See the findings

Safe, private and kind. We work only with written consent from the school, the teachers and the children's guardians, and a teacher can stop any session at any time. Children are identified by system IDs, access is role-based, and no personal details are sent to outside services. Child screens give positive, non-diagnostic feedback, and sessions are limited to 20 minutes with a calm break prompt.

You said, we improved

At our progress presentation the panel gave guidance, and we acted on it:

Sinhala & local accents

Added Sinhala audio and visuals, accent-tolerant speech checks, and UK-English pronunciation voices that suit Sri Lankan learners.

Backward navigation

Children and teachers can now step back between phases to review progress.

Beyond rules, carefully

We tested a learning model alongside the rules. It predicts pace better than a simple formula, but stays switched off until it is verified with more children.

Real classroom feedback

We ran sessions with children at special education schools and refined the app from teachers' observations.

Technologies used

Our toolbox: a tablet app, a secure server, specialist services for speech, handwriting and concepts, and cloud storage.

React Native + Expo

The tablet-first app used by children and teachers, built mainly for Android tablets.

Frontend

Node.js + Express

Secure API for sign-in, learning rules and reports, with role-based access and request limits.

Backend

Python (FastAPI, Flask)

Speech scoring (wav2vec2, Whisper), handwriting analysis (DTW, scikit-learn) and concept analysis.

Intelligence

PostgreSQL

The system of record for students, sessions, attempts and learning history.

Database

Neo4j

The concept knowledge graph that links students, concepts, categories and confusions.

Knowledge

Docker + cloud

Containerised services on managed cloud hosting, with a hosted language model only for anonymised teacher summaries.

Deployment

Honest research status. Not every model we explored is in the live system. The concept graph-attention model did not beat the simpler rules, so it stays in test mode. In every module, machine learning never decides a child's progress; clear rules and the teacher do.

Findings

These are the results reported in the four individual theses. Where a study is still running, we say so plainly.

Concept Learning: it works, and one idea did not

80.3%of concepts mastered at Tier 1 (98.3% at Tier 2)
+0.36average gain when a concept is practised again (p < 0.001)
441attempts from 19 children over 27 days
96.3%of untried concepts get a smart starting guess

Children improved reliably the more they practised a word. Concepts that look alike were the ones children mixed up (a clear, medium-sized effect), but concepts that merely sound alike were not. The graph-attention model did not beat simple rules, and the team reported that honestly. The study also found that the app's own word-choice policy muddies this kind of test, a useful lesson for future research. A formal usability score has not been measured yet.

Handwriting: transparent and reproducible

7learners in the controlled data collection
420letter-attempt records captured
7,609automated tests passed
0.41cluster separation score for two motor-pattern groups

The research code reproduced the app's own feature values exactly. Three different clustering methods found the same two groups, but the groups were sensitive to removing individual children, so they are used only to describe patterns, never to label a child. Score weights still need calibrating with more data and teacher ratings. Teacher usability results are being compiled.

Pronunciation: scores that agree with teachers

0.73agreement with teachers (correlation) on 32 reviewed attempts
7.9points average difference from the teacher's score
87.5%agreement on pass or fail
0.19agreement when using only simple word matching

Moving from word-level template matching to sound-by-sound scoring lifted agreement with teachers from 0.19 to 0.73, which shows why phoneme-level scoring matters for children's voices. These results are preliminary: they rest on 32 teacher-reviewed attempts out of 249 recorded (95% confidence range 0.51 to 0.86). The adaptation model for autistic children's speech is built and ready, and is waiting for teachers to finish labelling 167 classroom clips.

Dialogue: a better pace predictor, field study running

30everyday words taught in Level 1
114unit tests passing
0.70pace prediction accuracy (balanced) vs 0.50 for a plain formula
3schools in the supervised field study

On 87 real first-session records, a Random Forest predicted each child's learning pace clearly better than a literature-based formula for spoken attempts. The team also found and fixed early labelling faults by auditing rather than trusting headline accuracy. Results for speech scoring, echo detection and teacher usability will follow once the field study finishes.

What is still pending. Teacher usability scores (Concept and Handwriting), the Dialogue field-study results, and teacher labelling for the Pronunciation adaptation model. We will update this page as each result is confirmed.

Read the four theses

Curious to know more?

See how the project is progressing or meet the team behind it.

Document

Open in new tab Download