Rethinking Pronunciation Learning

Rethinking Pronunciation Learning

Rethinking Pronunciation Learning

If pronunciation learning is ultimately about speaking clearly and being understood, how should a digital learning system structure practice around that goal?

SayChinese is a research-driven, AI-supported Mandarin Chinese pronunciation platform that turns insights from language learning and speech research into a learner-controlled practice experience.

Designed, developed, and launched as a live public product.

MY ROLE

MY ROLE

MY ROLE

Researcher · Product Designer · Developer

Researcher · Product Designer · Developer

SayChinese began as a research-driven exploration of how digital pronunciation practice could better support connected speech. I led the project from concept and system design to AI-assisted development, public launch, and learner evaluation.

SayChinese began as a research-driven exploration of how digital pronunciation practice could better support connected speech. I led the project from concept and system design to AI-assisted development, public launch, and learner evaluation.

Research & Problem Framing

I translated linguistic, cognitive, and learning research into product principles that defined what the learning experience should support.

Product & System Design

I turned those principles into a multi-layer practice system, shaping the learning structure, interaction flow, and end-to-end experience.

AI-Assisted Development

I used AI-assisted coding and API integration to move from prototype to live product, while defining the segmentation logic, evaluating technical trade-offs, and retaining control over key product decisions.

Evaluation & Iteration

I studied how learners used the product, combining behavioural and qualitative evidence to identify patterns and guide the next product direction.

Research & Problem Framing

I translated linguistic, cognitive, and learning research into product principles that defined what the learning experience should support.

Product & System Design

I turned those principles into a multi-layer practice system, shaping the learning structure, interaction flow, and end-to-end experience.

AI-Assisted Development

I used AI-assisted coding and API integration to move from prototype to live product, while defining the segmentation logic, evaluating technical trade-offs, and retaining control over key product decisions.

Evaluation & Iteration

I studied how learners used the product, combining behavioural and qualitative evidence to identify patterns and guide the next product direction.

THE PROBLEM

THE PROBLEM

Individual Accuracy ≠ Connected Speech

Individual Accuracy ≠ Connected Speech

Individual Accuracy ≠ Connected Speech

In teaching Mandarin Chinese, I repeatedly saw learners recognise and pronounce individual words and tones accurately, yet struggle when those same words came together in a sentence.

In teaching Mandarin Chinese, I repeatedly saw learners recognise and pronounce individual words and tones accurately, yet struggle when those same words came together in a sentence.

Knowing the words was not always enough.

Knowing the words was not always enough.

In natural speech, words are not heard or produced as neat, separate units. Rhythm, phrasing, and meaningful chunks shape how a sentence is understood and spoken. Learners could know every part of a sentence and still struggle to hear how it was organised or make their own speech flow naturally.

In natural speech, words are not heard or produced as neat, separate units. Rhythm, phrasing, and meaningful chunks shape how a sentence is understood and spoken. Learners could know every part of a sentence and still struggle to hear how it was organised or make their own speech flow naturally.

This is especially challenging in Mandarin Chinese. Pitch can change a word’s meaning, and when words are spoken together, learners also need to manage how tones, rhythm, phrasing, and intonation work across the sentence.

This is especially challenging in Mandarin Chinese. Pitch can change a word’s meaning, and when words are spoken together, learners also need to manage how tones, rhythm, phrasing, and intonation work across the sentence.

This led to a design question:

This led to a design question:

How might a digital learning system support

both focused pronunciation practice and connected speech?

THE RESEARCH

THE RESEARCH

Designing Around How Speech Is Learned

Designing Around How Speech Is Learned

Designing Around How Speech Is Learned

The teaching problem pushed me to ask a more fundamental question: how do learners actually process and produce connected speech?

The teaching problem pushed me to ask a more fundamental question: how do learners actually process and produce connected speech?

I drew on linguistics, speech research, cognitive psychology, and language pedagogy. Three ideas shaped the design.

I drew on linguistics, speech research, cognitive psychology, and language pedagogy. Three ideas shaped the design.

1. Prosody gives speech its shape.

1. Prosody gives speech its shape.

Speech is not heard word by word. Rhythm, phrasing, pauses, tone, and intonation help us make sense of how it flows.

Speech is not heard word by word. Rhythm, phrasing, pauses, tone, and intonation help us make sense of how it flows.

>>> Design implication: Make phrase-level rhythm and grouping visible.

>>> Design implication: Make phrase-level rhythm and grouping visible.

2. Learners need manageable chunks.

2. Learners need manageable chunks.

Pronunciation asks learners to process several things at once. Smaller chunks make difficult details easier to focus on, while larger phrases preserve rhythm and flow.

Pronunciation asks learners to process several things at once. Smaller chunks make difficult details easier to focus on, while larger phrases preserve rhythm and flow.

>>> Design implication: Support more than one practice size.

>>> Design implication: Support more than one practice size.

3. Learners need control over how they practise.

3. Learners need control over how they practise.

What learners need can change during a single sentence — from checking one difficult word to returning to the whole phrase or sentence.

What learners need can change during a single sentence — from checking one difficult word to returning to the whole phrase or sentence.

>>> Design implication: Let learners move between levels as their needs change.

>>> Design implication: Let learners move between levels as their needs change.

Together, these insights pointed to a three-layer practice model: full-sentence audio, phrase-level practice, and smaller meaningful units, with learners free to move between levels as their needs changed.

Together, these insights pointed to a three-layer practice model: full-sentence audio, phrase-level practice, and smaller meaningful units, with learners free to move between levels as their needs changed.

THE DESIGN

THE DESIGN

One Sentence, Multiple Ways to Practise

One Sentence, Multiple Ways to Practise

The research pointed to a simple design principle: learners need different levels of support, but they should not be locked into one fixed path.

The research pointed to a simple design principle: learners need different levels of support, but they should not be locked into one fixed path.

I designed SayChinese as a three-layer practice system: full-sentence audio as a shared reference, with two switchable practice modes for phrase-level and smaller-unit practice.

I designed SayChinese as a three-layer practice system: full-sentence audio as a shared reference, with two switchable practice modes for phrase-level and smaller-unit practice.

Learners could switch freely between the two modes while keeping the full sentence available as a shared reference throughout practice.

Learners could switch freely between the two modes while keeping the full sentence available as a shared reference throughout practice.

DESIGNING THE PRACTICE UNITS

DESIGNING THE PRACTICE UNITS

The Logic Behind the Segmentation

The Logic Behind the Segmentation

The practice units in SayChinese were not created by simply splitting sentences into shorter pieces. I defined the segmentation logic myself, using my expertise in Mandarin Chinese linguistics, speech processing, and language teaching.

The practice units in SayChinese were not created by simply splitting sentences into shorter pieces. I defined the segmentation logic myself, using my expertise in Mandarin Chinese linguistics, speech processing, and language teaching.

The goal was to create units that were both meaningful and useful for practice.

The goal was to create units that were both meaningful and useful for practice.

Natural Speech Phrases

Natural Speech Phrases

I grouped words by meaning, sentence structure, rhythm, and natural pauses, creating larger units that preserve how the sentence flows in connected speech.

I grouped words by meaning, sentence structure, rhythm, and natural pauses, creating larger units that preserve how the sentence flows in connected speech.

Smaller Meaningful Units

Smaller Meaningful Units

I divided the same sentence into shorter, practiceable units while preserving meaningful words and chunks rather than splitting mechanically by character.

I divided the same sentence into shorter, practiceable units while preserving meaningful words and chunks rather than splitting mechanically by character.

The two switchable modes were designed to support different levels of attention within the same sentence.

Segmentation became part of the learning design,

not just a technical text-processing step.

Segmentation became part of the learning design,

not just a technical text-processing step.

BUILDING THE PRODUCT

AI Accelerated the Build, While the Learning Logic Stayed Human-Led

AI Accelerated the Build, While the Learning Logic Stayed Human-Led

AI Accelerated the Build, While the Learning Logic Stayed Human-Led

Once the learning architecture was defined, I used Claude Code in VS Code to move from design to a working public product. APIs supported pinyin generation, English translation, and audio, while the segmentation logic, interaction design, and learning decisions remained human-led.

Once the learning architecture was defined, I used Claude Code in VS Code to move from design to a working public product. APIs supported pinyin generation, English translation, and audio, while the segmentation logic, interaction design, and learning decisions remained human-led.

Rather than treating AI-generated code as finished work, I used an iterative build workflow that combined AI-assisted implementation with human review, verification, and product judgement.

Rather than treating AI-generated code as finished work, I used an iterative build workflow that combined AI-assisted implementation with human review, verification, and product judgement.

The build translated the three-layer learning model into a working interface: a shared full-sentence reference, two switchable practice modes, and learner-controlled movement between them.

The build translated the three-layer learning model into a working interface: a shared full-sentence reference, two switchable practice modes, and learner-controlled movement between them.

Human-led decisions defined the product;

AI-assisted execution accelerated how it was built.

Human-led decisions defined the product;

AI-assisted execution accelerated how it was built.

EVALUATING THE EXPERIENCE

EVALUATING THE EXPERIENCE

Learners Used the Layers as Complementary Support

Learners Used the Layers as Complementary Support

Learners Used the Layers as Complementary Support

Once the product was working, I wanted to understand how learners would actually use the three-layer system, not just which mode they preferred.

Once the product was working, I wanted to understand how learners would actually use the three-layer system, not just which mode they preferred.

I evaluated SayChinese with 16 adult Mandarin Chinese learners using behavioural logs, questionnaires, observation, and interviews. Rather than measuring pronunciation gains at this stage, I focused on how learners chose, switched, and combined the available support.

I evaluated SayChinese with 16 adult Mandarin Chinese learners using behavioural logs, questionnaires, observation, and interviews. Rather than measuring pronunciation gains at this stage, I focused on how learners chose, switched, and combined the available support.

1. Different Layers Supported Different Needs

1. Different Layers Supported Different Needs

Learners used the two modes for different kinds of practice. Natural Speech Phrases supported rhythm, phrasing, and connected speech, while Smaller Meaningful Units supported focused repetition, difficult words and tones, and targeted repair.

Learners used the two modes for different kinds of practice. Natural Speech Phrases supported rhythm, phrasing, and connected speech, while Smaller Meaningful Units supported focused repetition, difficult words and tones, and targeted repair.

2. Learners Moved Between Modes to Match Their Needs

2. Learners Moved Between Modes to Match Their Needs

Learners rarely treated one mode as a fixed destination. They moved between phrase-level and smaller-unit practice as their needs changed, zooming in for closer attention and back out to reconnect with rhythm and flow.

Learners rarely treated one mode as a fixed destination. They moved between phrase-level and smaller-unit practice as their needs changed, zooming in for closer attention and back out to reconnect with rhythm and flow.

3. The Full Sentence Kept the Layers Connected

3. The Full Sentence Kept the Layers Connected

Learners used full-sentence audio as a shared reference across both practice modes. They returned to it throughout practice to reconnect focused detail with the rhythm and flow of the complete utterance.

Learners used full-sentence audio as a shared reference across both practice modes. They returned to it throughout practice to reconnect focused detail with the rhythm and flow of the complete utterance.

4. Different Learners Needed Different Paths

4. Different Learners Needed Different Paths

The same three-layer system supported different practice strategies. Beginners relied more on listening, repetition, and smaller-unit support, while intermediate learners leaned more on phrase-level practice and used smaller units selectively when difficulties emerged.

The same three-layer system supported different practice strategies. Beginners relied more on listening, repetition, and smaller-unit support, while intermediate learners leaned more on phrase-level practice and used smaller units selectively when difficulties emerged.

One flexible system supported different practice strategies

without prescribing a single path.

One flexible system supported different practice strategies

without prescribing a single path.

DESIGN REFLECTION

DESIGN REFLECTION

The Evidence Showed That Choice Alone Wasn’t Enough

The Evidence Showed That Choice Alone Wasn’t Enough

I started with a simple principle:

give learners different levels of support and let them decide when to move between them.

I started with a simple principle:

give learners different levels of support and let them decide when to move between them.

Design Alone Isn't Enough

Design Alone Isn't Enough

The evaluation confirmed that flexibility mattered, but it also revealed the next design challenge. Learners were already switching levels for different reasons and at different moments.

The opportunity was no longer to add more modes,

but to help learners recognise when a different level of support might be useful

without taking that choice away from them.

The evaluation confirmed that flexibility mattered, but it also revealed the next design challenge. Learners were already switching levels for different reasons and at different moments.

The opportunity was no longer to add more modes,

but to help learners recognise when a different level of support might be useful

without taking that choice away from them.

The next design problem became guidance, not more options.

The next design problem became guidance, not more options.

NEXT ITERATION

NEXT ITERATION

From Flexible Practice to More Guided Support

From Flexible Practice to More Guided Support

From Flexible Practice to More Guided Support

The next version would build on the existing three-layer system rather than replace it.

The next version would build on the existing three-layer system rather than replace it.

Recording for self-awareness

Recording for self-awareness

Let learners record their own speech and revisit attempts, making production more visible within the practice loop.

Let learners record their own speech and revisit attempts, making production more visible within the practice loop.

More informative feedback

More informative feedback

Move beyond playback toward feedback that helps learners notice where pronunciation, rhythm, or connected speech may need more attention.

Move beyond playback toward feedback that helps learners notice where pronunciation, rhythm, or connected speech may need more attention.

Conversational AI for practice in context

Conversational AI for practice in context

Explore short AI-supported exchanges where learners can apply pronunciation practice in more communicative, sentence-level interaction.

Explore short AI-supported exchanges where learners can apply pronunciation practice in more communicative, sentence-level interaction.

The goal is not to automate the learner’s path, but to provide better signals for

what to practise next while keeping control with the learner.

The goal is not to automate the learner’s path, but to provide better signals for what to practise next while keeping control with the learner.

The goal is not to automate the learner’s path, but to provide better signals for

what to practise next while keeping control with the learner.

CURRENT STAGE

CURRENT STAGE

The Live Product Is Still Informing the Next Version

The Live Product Is Still Informing the Next Version

The Live Product Is Still Informing the Next Version

SayChinese is now publicly available and continuing to collect feedback from a broader learner base.

SayChinese is now publicly available and continuing to collect feedback from a broader learner base.

Rather than treating the roadmap as fixed, I am using the public release to understand how the product performs beyond the initial study: which features learners return to, where friction appears, and what support they ask for next.

Rather than treating the roadmap as fixed, I am using the public release to understand how the product performs beyond the initial study: which features learners return to, where friction appears, and what support they ask for next.

Those signals will help determine which direction recording, richer feedback, or conversational practice should be prioritised in the next iteration.

Those signals will help determine which direction recording, richer feedback, or conversational practice should be prioritised in the next iteration.