Rethinking Pronunciation Learning
Rethinking Pronunciation Learning
Rethinking Pronunciation Learning
If pronunciation learning is ultimately about speaking clearly and being understood, how should a digital learning system structure practice around that goal?
SayChinese is a research-driven, AI-supported Mandarin Chinese pronunciation platform that turns insights from language learning and speech research into a learner-controlled practice experience.
Designed, developed, and launched as a live public product.
MY ROLE
MY ROLE
MY ROLE
Researcher · Product Designer · Developer
Researcher · Product Designer · Developer
SayChinese began as a research-driven exploration of how digital pronunciation practice could better support connected speech. I led the project from concept and system design to AI-assisted development, public launch, and learner evaluation.
SayChinese began as a research-driven exploration of how digital pronunciation practice could better support connected speech. I led the project from concept and system design to AI-assisted development, public launch, and learner evaluation.
Research & Problem Framing
I translated linguistic, cognitive, and learning research into product principles that defined what the learning experience should support.
Product & System Design
I turned those principles into a multi-layer practice system, shaping the learning structure, interaction flow, and end-to-end experience.
AI-Assisted Development
I used AI-assisted coding and API integration to move from prototype to live product, while defining the segmentation logic, evaluating technical trade-offs, and retaining control over key product decisions.
Evaluation & Iteration
I studied how learners used the product, combining behavioural and qualitative evidence to identify patterns and guide the next product direction.
Research & Problem Framing
I translated linguistic, cognitive, and learning research into product principles that defined what the learning experience should support.
Product & System Design
I turned those principles into a multi-layer practice system, shaping the learning structure, interaction flow, and end-to-end experience.
AI-Assisted Development
I used AI-assisted coding and API integration to move from prototype to live product, while defining the segmentation logic, evaluating technical trade-offs, and retaining control over key product decisions.
Evaluation & Iteration
I studied how learners used the product, combining behavioural and qualitative evidence to identify patterns and guide the next product direction.
THE PROBLEM
THE PROBLEM
Individual Accuracy ≠ Connected Speech
Individual Accuracy ≠ Connected Speech
Individual Accuracy ≠ Connected Speech
In teaching Mandarin Chinese, I repeatedly saw learners recognise and pronounce individual words and tones accurately, yet struggle when those same words came together in a sentence.
In teaching Mandarin Chinese, I repeatedly saw learners recognise and pronounce individual words and tones accurately, yet struggle when those same words came together in a sentence.
Knowing the words was not always enough.
Knowing the words was not always enough.
In natural speech, words are not heard or produced as neat, separate units. Rhythm, phrasing, and meaningful chunks shape how a sentence is understood and spoken. Learners could know every part of a sentence and still struggle to hear how it was organised or make their own speech flow naturally.
In natural speech, words are not heard or produced as neat, separate units. Rhythm, phrasing, and meaningful chunks shape how a sentence is understood and spoken. Learners could know every part of a sentence and still struggle to hear how it was organised or make their own speech flow naturally.
This is especially challenging in Mandarin Chinese. Pitch can change a word’s meaning, and when words are spoken together, learners also need to manage how tones, rhythm, phrasing, and intonation work across the sentence.
This is especially challenging in Mandarin Chinese. Pitch can change a word’s meaning, and when words are spoken together, learners also need to manage how tones, rhythm, phrasing, and intonation work across the sentence.
This led to a design question:
This led to a design question:
How might a digital learning system support
both focused pronunciation practice and connected speech?
THE RESEARCH
THE RESEARCH
Designing Around How Speech Is Learned
Designing Around How Speech Is Learned
Designing Around How Speech Is Learned
The teaching problem pushed me to ask a more fundamental question: how do learners actually process and produce connected speech?
The teaching problem pushed me to ask a more fundamental question: how do learners actually process and produce connected speech?
I drew on linguistics, speech research, cognitive psychology, and language pedagogy. Three ideas shaped the design.
I drew on linguistics, speech research, cognitive psychology, and language pedagogy. Three ideas shaped the design.

1. Prosody gives speech its shape.
1. Prosody gives speech its shape.
Speech is not heard word by word. Rhythm, phrasing, pauses, tone, and intonation help us make sense of how it flows.
Speech is not heard word by word. Rhythm, phrasing, pauses, tone, and intonation help us make sense of how it flows.
>>> Design implication: Make phrase-level rhythm and grouping visible.
>>> Design implication: Make phrase-level rhythm and grouping visible.

2. Learners need manageable chunks.
2. Learners need manageable chunks.
Pronunciation asks learners to process several things at once. Smaller chunks make difficult details easier to focus on, while larger phrases preserve rhythm and flow.
Pronunciation asks learners to process several things at once. Smaller chunks make difficult details easier to focus on, while larger phrases preserve rhythm and flow.
>>> Design implication: Support more than one practice size.
>>> Design implication: Support more than one practice size.

3. Learners need control over how they practise.
3. Learners need control over how they practise.
What learners need can change during a single sentence — from checking one difficult word to returning to the whole phrase or sentence.
What learners need can change during a single sentence — from checking one difficult word to returning to the whole phrase or sentence.
>>> Design implication: Let learners move between levels as their needs change.
>>> Design implication: Let learners move between levels as their needs change.
Together, these insights pointed to a three-layer practice model: full-sentence audio, phrase-level practice, and smaller meaningful units, with learners free to move between levels as their needs changed.
Together, these insights pointed to a three-layer practice model: full-sentence audio, phrase-level practice, and smaller meaningful units, with learners free to move between levels as their needs changed.
THE DESIGN
THE DESIGN
One Sentence, Multiple Ways to Practise
One Sentence, Multiple Ways to Practise
The research pointed to a simple design principle: learners need different levels of support, but they should not be locked into one fixed path.
The research pointed to a simple design principle: learners need different levels of support, but they should not be locked into one fixed path.
I designed SayChinese as a three-layer practice system: full-sentence audio as a shared reference, with two switchable practice modes for phrase-level and smaller-unit practice.
I designed SayChinese as a three-layer practice system: full-sentence audio as a shared reference, with two switchable practice modes for phrase-level and smaller-unit practice.

Learners could switch freely between the two modes while keeping the full sentence available as a shared reference throughout practice.
Learners could switch freely between the two modes while keeping the full sentence available as a shared reference throughout practice.
DESIGNING THE PRACTICE UNITS
DESIGNING THE PRACTICE UNITS
The Logic Behind the Segmentation
The Logic Behind the Segmentation
The practice units in SayChinese were not created by simply splitting sentences into shorter pieces. I defined the segmentation logic myself, using my expertise in Mandarin Chinese linguistics, speech processing, and language teaching.
The practice units in SayChinese were not created by simply splitting sentences into shorter pieces. I defined the segmentation logic myself, using my expertise in Mandarin Chinese linguistics, speech processing, and language teaching.
The goal was to create units that were both meaningful and useful for practice.
The goal was to create units that were both meaningful and useful for practice.
Natural Speech Phrases
Natural Speech Phrases
I grouped words by meaning, sentence structure, rhythm, and natural pauses, creating larger units that preserve how the sentence flows in connected speech.
I grouped words by meaning, sentence structure, rhythm, and natural pauses, creating larger units that preserve how the sentence flows in connected speech.
Smaller Meaningful Units
Smaller Meaningful Units
I divided the same sentence into shorter, practiceable units while preserving meaningful words and chunks rather than splitting mechanically by character.
I divided the same sentence into shorter, practiceable units while preserving meaningful words and chunks rather than splitting mechanically by character.
The two switchable modes were designed to support different levels of attention within the same sentence.

Segmentation became part of the learning design,
not just a technical text-processing step.
Segmentation became part of the learning design,
not just a technical text-processing step.
BUILDING THE PRODUCT
AI Accelerated the Build, While the Learning Logic Stayed Human-Led
AI Accelerated the Build, While the Learning Logic Stayed Human-Led
AI Accelerated the Build, While the Learning Logic Stayed Human-Led
Once the learning architecture was defined, I used Claude Code in VS Code to move from design to a working public product. APIs supported pinyin generation, English translation, and audio, while the segmentation logic, interaction design, and learning decisions remained human-led.
Once the learning architecture was defined, I used Claude Code in VS Code to move from design to a working public product. APIs supported pinyin generation, English translation, and audio, while the segmentation logic, interaction design, and learning decisions remained human-led.
Rather than treating AI-generated code as finished work, I used an iterative build workflow that combined AI-assisted implementation with human review, verification, and product judgement.
Rather than treating AI-generated code as finished work, I used an iterative build workflow that combined AI-assisted implementation with human review, verification, and product judgement.
The build translated the three-layer learning model into a working interface: a shared full-sentence reference, two switchable practice modes, and learner-controlled movement between them.
The build translated the three-layer learning model into a working interface: a shared full-sentence reference, two switchable practice modes, and learner-controlled movement between them.

Human-led decisions defined the product;
AI-assisted execution accelerated how it was built.
Human-led decisions defined the product;
AI-assisted execution accelerated how it was built.
EVALUATING THE EXPERIENCE
EVALUATING THE EXPERIENCE
Learners Used the Layers as Complementary Support
Learners Used the Layers as Complementary Support
Learners Used the Layers as Complementary Support
Once the product was working, I wanted to understand how learners would actually use the three-layer system, not just which mode they preferred.
Once the product was working, I wanted to understand how learners would actually use the three-layer system, not just which mode they preferred.
I evaluated SayChinese with 16 adult Mandarin Chinese learners using behavioural logs, questionnaires, observation, and interviews. Rather than measuring pronunciation gains at this stage, I focused on how learners chose, switched, and combined the available support.
I evaluated SayChinese with 16 adult Mandarin Chinese learners using behavioural logs, questionnaires, observation, and interviews. Rather than measuring pronunciation gains at this stage, I focused on how learners chose, switched, and combined the available support.

1. Different Layers Supported Different Needs
1. Different Layers Supported Different Needs
Learners used the two modes for different kinds of practice. Natural Speech Phrases supported rhythm, phrasing, and connected speech, while Smaller Meaningful Units supported focused repetition, difficult words and tones, and targeted repair.
Learners used the two modes for different kinds of practice. Natural Speech Phrases supported rhythm, phrasing, and connected speech, while Smaller Meaningful Units supported focused repetition, difficult words and tones, and targeted repair.

2. Learners Moved Between Modes to Match Their Needs
2. Learners Moved Between Modes to Match Their Needs
Learners rarely treated one mode as a fixed destination. They moved between phrase-level and smaller-unit practice as their needs changed, zooming in for closer attention and back out to reconnect with rhythm and flow.
Learners rarely treated one mode as a fixed destination. They moved between phrase-level and smaller-unit practice as their needs changed, zooming in for closer attention and back out to reconnect with rhythm and flow.

3. The Full Sentence Kept the Layers Connected
3. The Full Sentence Kept the Layers Connected
Learners used full-sentence audio as a shared reference across both practice modes. They returned to it throughout practice to reconnect focused detail with the rhythm and flow of the complete utterance.
Learners used full-sentence audio as a shared reference across both practice modes. They returned to it throughout practice to reconnect focused detail with the rhythm and flow of the complete utterance.

4. Different Learners Needed Different Paths
4. Different Learners Needed Different Paths
The same three-layer system supported different practice strategies. Beginners relied more on listening, repetition, and smaller-unit support, while intermediate learners leaned more on phrase-level practice and used smaller units selectively when difficulties emerged.
The same three-layer system supported different practice strategies. Beginners relied more on listening, repetition, and smaller-unit support, while intermediate learners leaned more on phrase-level practice and used smaller units selectively when difficulties emerged.

One flexible system supported different practice strategies
without prescribing a single path.
One flexible system supported different practice strategies
without prescribing a single path.
DESIGN REFLECTION
DESIGN REFLECTION
The Evidence Showed That Choice Alone Wasn’t Enough
The Evidence Showed That Choice Alone Wasn’t Enough
I started with a simple principle:
give learners different levels of support and let them decide when to move between them.
I started with a simple principle:
give learners different levels of support and let them decide when to move between them.
Design Alone Isn't Enough
Design Alone Isn't Enough
The evaluation confirmed that flexibility mattered, but it also revealed the next design challenge. Learners were already switching levels for different reasons and at different moments.
The opportunity was no longer to add more modes,
but to help learners recognise when a different level of support might be useful
without taking that choice away from them.
The evaluation confirmed that flexibility mattered, but it also revealed the next design challenge. Learners were already switching levels for different reasons and at different moments.
The opportunity was no longer to add more modes,
but to help learners recognise when a different level of support might be useful
without taking that choice away from them.
The next design problem became guidance, not more options.
The next design problem became guidance, not more options.
NEXT ITERATION
NEXT ITERATION
From Flexible Practice to More Guided Support
From Flexible Practice to More Guided Support
From Flexible Practice to More Guided Support
The next version would build on the existing three-layer system rather than replace it.
The next version would build on the existing three-layer system rather than replace it.
Recording for self-awareness
Recording for self-awareness
Let learners record their own speech and revisit attempts, making production more visible within the practice loop.
Let learners record their own speech and revisit attempts, making production more visible within the practice loop.
More informative feedback
More informative feedback
Move beyond playback toward feedback that helps learners notice where pronunciation, rhythm, or connected speech may need more attention.
Move beyond playback toward feedback that helps learners notice where pronunciation, rhythm, or connected speech may need more attention.
Conversational AI for practice in context
Conversational AI for practice in context
Explore short AI-supported exchanges where learners can apply pronunciation practice in more communicative, sentence-level interaction.
Explore short AI-supported exchanges where learners can apply pronunciation practice in more communicative, sentence-level interaction.
The goal is not to automate the learner’s path, but to provide better signals for
what to practise next while keeping control with the learner.
The goal is not to automate the learner’s path, but to provide better signals for what to practise next while keeping control with the learner.
The goal is not to automate the learner’s path, but to provide better signals for
what to practise next while keeping control with the learner.
CURRENT STAGE
CURRENT STAGE
The Live Product Is Still Informing the Next Version
The Live Product Is Still Informing the Next Version
The Live Product Is Still Informing the Next Version
SayChinese is now publicly available and continuing to collect feedback from a broader learner base.
SayChinese is now publicly available and continuing to collect feedback from a broader learner base.
Rather than treating the roadmap as fixed, I am using the public release to understand how the product performs beyond the initial study: which features learners return to, where friction appears, and what support they ask for next.
Rather than treating the roadmap as fixed, I am using the public release to understand how the product performs beyond the initial study: which features learners return to, where friction appears, and what support they ask for next.
Those signals will help determine which direction recording, richer feedback, or conversational practice should be prioritised in the next iteration.
Those signals will help determine which direction recording, richer feedback, or conversational practice should be prioritised in the next iteration.