D
Dashi Dance
Guest
My wife is a dancer, and for years I watched her fight with her tools instead of practicing. Learning a routine from a video usually means mirroring it, slowing it down, and drilling one section over and over. A normal video player isn't really built for that. There's no A/B loop, so you drag the playhead back by hand every rep. The speed steps jump from 0.5x straight to 1x. Comparing your run with the original means opening a video editor and spending a lot of time getting things lined up. And if you want to record yourself while the reference is playing, you need two devices.
So I built the app she wished existed.
Dashi Dance went live on the App Store on August 25, for iPhone and iPad, in 15 languages. The first commit, a requirements document, landed on June 16, so it took about ten weeks.
The interesting part is that when I started, I had written zero Swift and knew nothing about iOS. I'm an experienced software engineer, but iOS was completely new to me. From the first day, I built the app with Claude Code, Anthropic's AI coding agent. It was basically a team of two: me and the agent. I honestly wasn't sure that would work.
This is what that looked like, including the parts that didn't go well.
Everything in Dashi Dance is one loop, and a dancer runs that loop hundreds of times on a single routine:
Then you go back to step two.
That whole loop is free, supported by ads. Plus removes the ads and adds themes; Pro adds things like iCloud sync, EQ and pitch controls for the reference, and automatic alignment based on the audio.
The first thing I'd tell anyone trying this is that getting an agent to write code quickly isn't the hard part. The hard part is getting it to stop writing the wrong code quickly and confidently. You need a process around it.
So almost nothing was built from a one-line prompt. Features started as a spec, went through a back-and-forth until I agreed with it, then became a plan broken into tasks. The implementation was reviewed against the plan afterward.
As I write this, the repository has more than 90 specs, more than 110 plans, and over 4,400 commits. There's CI, unit tests, UI tests, and a pipeline for localization and App Store publishing. None of that is particularly glamorous, but it turned out to be necessary. The agent doesn't remember yesterday, so the repository has to.
The single most useful habit we developed was turning repeated mistakes into mechanisms instead of adding more rules. Early on, whenever the agent did something wrong, I'd add a line to its instructions file. That file grew past a thousand lines, and the same mistakes kept happening anyway. A rule only helps if it gets read and remembered at the right moment.
So we started asking a different question: what would make this mistake impossible?
The answers became lint rules, test gates, and a hook that checks every shell command the agent runs before it executes. If a command is one we've had problems with before, the hook refuses it and tells the agent what to use instead.
The second habit was refusing claims without evidence. An agent will tell you a layout "looks correct" after reading a screenshot, and it can still be wrong about spacing or reachability in ways you notice immediately on a real phone. So if the agent makes a claim about where something sits on screen, I want numbers: frames, pixel measurements, or a diff. If it says a fix works, I want a test that fails without the fix.
It sounds pedantic. It has saved me more than once.
Once, nine separate code reviews of the same change missed a regression in a UI test. A reviewer reading a diff simply can't see everything a UI test will do. Now, if a view is touched, its UI tests run. No exceptions.
Comparison plays your take and the reference side by side, and none of it works unless they agree on time. Two independent video players drift apart over the length of a clip, so comparison runs both from a single clock derived from the system's host time.
The harder problem was the offset we create ourselves. An app has no way to start a recording at exactly the same instant playback starts. So the recording deliberately starts early and captures everything. That leaves your take a few hundred milliseconds longer than the music you actually danced to. We measured 0.399 and 0.450 seconds on two takes from one phone.
The app now measures that gap when recording starts and seeds it directly into the alignment. An in-app recording can therefore open already lined up.
On a real iPhone, the first tap on Play took 0.44 seconds to move the playhead, while the second took 0.03 seconds.
We went through five wrong theories, and each one looked reasonable enough that I was sure it was the one: the camera session, Apple's play-immediately API, pre-rolling the decoder, our own audio effects chain building lazily, and a warm-up for that chain that we shipped and reverted a day later.
What finally helped was one simple experiment: two copies of the same clip, one with the audio track removed. The silent one started in 60 to 234 milliseconds. The one with sound took 442 to 698 milliseconds, every trial. The difference turned out to be iOS starting its audio output pipeline. None of our code was on that path. So we accounted for it in the alignment and stopped chasing it. It's still a little annoying that the fix for a bug was deciding it wasn't a bug.
The iOS simulator is fast and forgiving, and it can make you think the app is finished.
Five bugs only ever showed up on a device: an iPad that froze coming back into the app after the system text size changed (its tab bar kept rebuilding its icons and never settled, and that one went to Apple as a bug report); a zoomed video pane that stole the other pane's touches; a fade effect sitting on top of a Share button, visible only on a phone with a taller safe area; a sticker viewer that brought back the previous sticker if you tapped quickly; and a crash when tapping a reminder notification with the app fully closed, a path no automated test could reach.
Some of the work was in fields I knew nothing about, and this is probably where the agent changed the project the most.
Each of the app's six color themes has two accent colors, and my whole brief was that the two shouldn't look like one. I started by measuring WCAG contrast, the accessibility ratio, and by that number our worst pair was a violet next to an oxblood at 1.02:1, practically identical.
Except anyone can see they're two colors. I didn't even know that was a different question.
Contrast ratio only measures luminance. The question "do these look the same?" is about perceptual distance, which has its own metric, CIEDE2000. Re-measured that way, the real offenders were two different themes in dark mode.
The fix moved lightness and left hue alone, and a test now fails the build if any pair drops below the floor. I learned all of that in a few days by having the agent explain each step and pushing back until it made sense.
Localization went much the same way. The app shipped in 15 languages with no translation agency. An LLM translates everything from scripts next to the code, and a termbase keeps named controls consistent: "A/B Loop" is translated once per language and handed to every later prompt.
Automated checks catch broken placeholders, leftover English, and help pages that don't use the app's own button labels. And we still read the output, because things that pass every check can still be wrong.
French shipped "Chorée" for "Choreo", which is chorea, a neurological disorder. Simplified Chinese used the storage word for "save" across 17 strings while the button itself said something else. A Japanese help page printed a whole section twice, with every gate green.
I read English and Chinese. For the other twelve languages, we rely on the checks, a second AI model that flags translations whose meaning has drifted from the source, and hopefully feedback from a dancer who speaks one of those languages when a word sounds wrong. That part still makes me a bit nervous.
Accessibility got an audit before launch, because the App Store lets you declare accessibility support and we claim seven features: VoiceOver, Voice Control, Larger Text, Dark Interface, Differentiate Without Color Alone, Sufficient Contrast, and Reduced Motion.
The audit found twelve buttons in seven files drawing white text directly onto the theme color. In dark mode, the contrast measured as low as 1.2:1. It also found a text field that VoiceOver couldn't name at all.
Nobody had reported either problem. I suspect someone who can't read a button just stops using the app rather than writing in.
There's no account and no server of ours. Stats and notifications live on the phone; sync goes through your private iCloud.
The place privacy cost us the most thinking was Photos. The app saves exports into its own albums, and the lazy way to do that is to ask for full library access. We dug into the smallest grant that still lets an app create an album and landed on Limited Access with zero photos selected. That lets the app file its exports while still preventing it from seeing anything already in your library. Importing needs no permission at all, since the system picker hands the app only the video you chose.
Even the bug report works this way. Contact in the app opens Mail with your message, a short block of diagnostics and, if you leave the switches on, a small on-device log and the latest crash report. Nothing leaves the phone until you press Send.
Since August 13, I've posted a build journal on X on most days, tagged for Shipaton, and it changed how I work more than I expected.
Writing a post means checking every claim against the code before it goes out. An early draft said the player had "four things to tap," based on a screenshot whose control bar actually had eight.
The honest part is that the audience is small. I shared each post into the Shipaton community for weeks without a reply, which was a little deflating. The journal still turned out useful as a record, and it's most of what this article was written from.
The feedback that did change something came from Reddit, where I introduced the app in a few dance communities. Several people couldn't see how an instructor, rather than a student, would get anything out of it. The app had an answer, but nothing on the website explained it clearly, so we added a Who It's For page that splits the use cases by role, from the dancer drilling between classes to the instructor teaching the routine.
Three updates shipped in the first two and a half weeks after launch: a help system with a "?" on every main screen that explains each control, a welcome tour with a sample project, a one-thumb camera zoom, and Dance Notes, timed text that you can pin onto the video itself, like "arm up on 8". The next release is already being built.
I don't have a neat summary yet. The closest I've got is that the agent is faster than I'll ever be, and the discipline has to come from the process, which is why so much of it is written down in the repo instead of in either of our heads.
If you learn dance from video, I'd love to hear what the most annoying part of your setup is. That's the question this whole app started from.
So I built the app she wished existed.
Dashi Dance went live on the App Store on August 25, for iPhone and iPad, in 15 languages. The first commit, a requirements document, landed on June 16, so it took about ten weeks.
The interesting part is that when I started, I had written zero Swift and knew nothing about iOS. I'm an experienced software engineer, but iOS was completely new to me. From the first day, I built the app with Claude Code, Anthropic's AI coding agent. It was basically a team of two: me and the agent. I honestly wasn't sure that would work.
This is what that looked like, including the parts that didn't go well.
What the app does
Everything in Dashi Dance is one loop, and a dancer runs that loop hundreds of times on a single routine:
- Import a reference video or song into a project.
- Learn it in the player: loop the eight counts that keep falling apart, slow them down with the music still on pitch, and mirror the video so left and right match your body.
- Record yourself with the reference playing beside the camera, with a countdown so you never have to touch the phone.
- Compare your take with the reference, side by side on the same clock.
Then you go back to step two.
That whole loop is free, supported by ads. Plus removes the ads and adds themes; Pro adds things like iCloud sync, EQ and pitch controls for the reference, and automatic alignment based on the audio.
Engineering, not vibe coding
The first thing I'd tell anyone trying this is that getting an agent to write code quickly isn't the hard part. The hard part is getting it to stop writing the wrong code quickly and confidently. You need a process around it.
So almost nothing was built from a one-line prompt. Features started as a spec, went through a back-and-forth until I agreed with it, then became a plan broken into tasks. The implementation was reviewed against the plan afterward.
As I write this, the repository has more than 90 specs, more than 110 plans, and over 4,400 commits. There's CI, unit tests, UI tests, and a pipeline for localization and App Store publishing. None of that is particularly glamorous, but it turned out to be necessary. The agent doesn't remember yesterday, so the repository has to.
The single most useful habit we developed was turning repeated mistakes into mechanisms instead of adding more rules. Early on, whenever the agent did something wrong, I'd add a line to its instructions file. That file grew past a thousand lines, and the same mistakes kept happening anyway. A rule only helps if it gets read and remembered at the right moment.
So we started asking a different question: what would make this mistake impossible?
The answers became lint rules, test gates, and a hook that checks every shell command the agent runs before it executes. If a command is one we've had problems with before, the hook refuses it and tells the agent what to use instead.
The second habit was refusing claims without evidence. An agent will tell you a layout "looks correct" after reading a screenshot, and it can still be wrong about spacing or reachability in ways you notice immediately on a real phone. So if the agent makes a claim about where something sits on screen, I want numbers: frames, pixel measurements, or a diff. If it says a fix works, I want a test that fails without the fix.
It sounds pedantic. It has saved me more than once.
Once, nine separate code reviews of the same change missed a regression in a UI test. A reviewer reading a diff simply can't see everything a UI test will do. Now, if a view is touched, its UI tests run. No exceptions.
The bugs that taught us the most
One clock for two videos
Comparison plays your take and the reference side by side, and none of it works unless they agree on time. Two independent video players drift apart over the length of a clip, so comparison runs both from a single clock derived from the system's host time.
The harder problem was the offset we create ourselves. An app has no way to start a recording at exactly the same instant playback starts. So the recording deliberately starts early and captures everything. That leaves your take a few hundred milliseconds longer than the music you actually danced to. We measured 0.399 and 0.450 seconds on two takes from one phone.
The app now measures that gap when recording starts and seeds it directly into the alignment. An in-app recording can therefore open already lined up.
The 420-millisecond ghost
On a real iPhone, the first tap on Play took 0.44 seconds to move the playhead, while the second took 0.03 seconds.
We went through five wrong theories, and each one looked reasonable enough that I was sure it was the one: the camera session, Apple's play-immediately API, pre-rolling the decoder, our own audio effects chain building lazily, and a warm-up for that chain that we shipped and reverted a day later.
What finally helped was one simple experiment: two copies of the same clip, one with the audio track removed. The silent one started in 60 to 234 milliseconds. The one with sound took 442 to 698 milliseconds, every trial. The difference turned out to be iOS starting its audio output pipeline. None of our code was on that path. So we accounted for it in the alignment and stopped chasing it. It's still a little annoying that the fix for a bug was deciding it wasn't a bug.
The simulator lies by omission
The iOS simulator is fast and forgiving, and it can make you think the app is finished.
Five bugs only ever showed up on a device: an iPad that froze coming back into the app after the system text size changed (its tab bar kept rebuilding its icons and never settled, and that one went to Apple as a bug report); a zoomed video pane that stole the other pane's touches; a fade effect sitting on top of a Share button, visible only on a phone with a taller safe area; a sticker viewer that brought back the previous sticker if you tapped quickly; and a crash when tapping a reminder notification with the app fully closed, a path no automated test could reach.
Learning a field from zero
Some of the work was in fields I knew nothing about, and this is probably where the agent changed the project the most.
Each of the app's six color themes has two accent colors, and my whole brief was that the two shouldn't look like one. I started by measuring WCAG contrast, the accessibility ratio, and by that number our worst pair was a violet next to an oxblood at 1.02:1, practically identical.
Except anyone can see they're two colors. I didn't even know that was a different question.
Contrast ratio only measures luminance. The question "do these look the same?" is about perceptual distance, which has its own metric, CIEDE2000. Re-measured that way, the real offenders were two different themes in dark mode.
The fix moved lightness and left hue alone, and a test now fails the build if any pair drops below the floor. I learned all of that in a few days by having the agent explain each step and pushing back until it made sense.
Localization went much the same way. The app shipped in 15 languages with no translation agency. An LLM translates everything from scripts next to the code, and a termbase keeps named controls consistent: "A/B Loop" is translated once per language and handed to every later prompt.
Automated checks catch broken placeholders, leftover English, and help pages that don't use the app's own button labels. And we still read the output, because things that pass every check can still be wrong.
French shipped "Chorée" for "Choreo", which is chorea, a neurological disorder. Simplified Chinese used the storage word for "save" across 17 strings while the button itself said something else. A Japanese help page printed a whole section twice, with every gate green.
I read English and Chinese. For the other twelve languages, we rely on the checks, a second AI model that flags translations whose meaning has drifted from the source, and hopefully feedback from a dancer who speaks one of those languages when a word sounds wrong. That part still makes me a bit nervous.
Accessibility got an audit before launch, because the App Store lets you declare accessibility support and we claim seven features: VoiceOver, Voice Control, Larger Text, Dark Interface, Differentiate Without Color Alone, Sufficient Contrast, and Reduced Motion.
The audit found twelve buttons in seven files drawing white text directly onto the theme color. In dark mode, the contrast measured as low as 1.2:1. It also found a text field that VoiceOver couldn't name at all.
Nobody had reported either problem. I suspect someone who can't read a button just stops using the app rather than writing in.
Privacy as a design input
There's no account and no server of ours. Stats and notifications live on the phone; sync goes through your private iCloud.
The place privacy cost us the most thinking was Photos. The app saves exports into its own albums, and the lazy way to do that is to ask for full library access. We dug into the smallest grant that still lets an app create an album and landed on Limited Access with zero photos selected. That lets the app file its exports while still preventing it from seeing anything already in your library. Importing needs no permission at all, since the system picker hands the app only the video you chose.
Even the bug report works this way. Contact in the app opens Mail with your message, a short block of diagnostics and, if you leave the switches on, a small on-device log and the latest crash report. Nothing leaves the phone until you press Send.
Building in public
Since August 13, I've posted a build journal on X on most days, tagged for Shipaton, and it changed how I work more than I expected.
Writing a post means checking every claim against the code before it goes out. An early draft said the player had "four things to tap," based on a screenshot whose control bar actually had eight.
The honest part is that the audience is small. I shared each post into the Shipaton community for weeks without a reply, which was a little deflating. The journal still turned out useful as a record, and it's most of what this article was written from.
The feedback that did change something came from Reddit, where I introduced the app in a few dance communities. Several people couldn't see how an instructor, rather than a student, would get anything out of it. The app had an answer, but nothing on the website explained it clearly, so we added a Who It's For page that splits the use cases by role, from the dancer drilling between classes to the instructor teaching the routine.
After launch
Three updates shipped in the first two and a half weeks after launch: a help system with a "?" on every main screen that explains each control, a welcome tour with a sample project, a one-thumb camera zoom, and Dance Notes, timed text that you can pin onto the video itself, like "arm up on 8". The next release is already being built.
I don't have a neat summary yet. The closest I've got is that the agent is faster than I'll ever be, and the discipline has to come from the process, which is why so much of it is written down in the repo instead of in either of our heads.
If you learn dance from video, I'd love to hear what the most annoying part of your setup is. That's the question this whole app started from.