Washa Guide

Washa Guide

Beta Test Edition

Document version
0.2.1
Issue date
2026-10-07
Target app
Washa (話者) app 0.2.0 (as of build 14, the serial number shared by all OSes)
Target OSes
iPhone (iOS), Mac (macOS), Android, Windows
Start reading
Washa on Mac: the recording list and a transcript
Washa on iPhone: the recording list
Washa on Apple Watch: a remote that stops the iPhone recording from your wrist

Teleport Inc. (テレポート株式会社)

1. What Washa is

1.1 In a nutshell

Washa (話者) is an app that turns audio or video recordings into text split into "who spoke, when, and what". Record a meeting, interview, class or discussion, or import an existing recording file, and it does the following entirely on the device:

The finished transcript can be read with a speaker and timestamp attached to each line. Once you give a speaker a name, Washa remembers that voice and, from the next recording on, suggests "Is this this person?"

Screen example: In a nutshell

1.2 Key features

Feature Details
Processes on the device Transcription, voiceprint analysis and speaker identification run inside the device you are using (iPhone, Mac, Android, Windows). Recordings are not sent to an AI company's servers.
Distinguishes up to 10 people Distinguishes up to 10 speakers in a single recording.
Japanese and English mixed is fine Choose "Auto (Japanese with some English)" and conversations in Japanese with English mixed in are transcribed as they are.
Multilingual Supports Japanese, English, Chinese (Mandarin), Cantonese and Korean. On iPhone and Mac you can also choose German, French, Spanish, Portuguese (Brazil) and Italian. Apart from the Japanese–English combination, each recording is treated as one language.
Free and unlimited at its core Recording, separating speakers, and using the text of who said what is free. There is no monthly fee and no cap on hours (Chapter 3).
Gets smarter with a word from a person When a person tells it "how many people were talking" or "whose voice this is", the result improves accordingly. Voices you teach it stay in the Speaker library and help in the next recording.
4 OSes Usable on iPhone, Mac, Android and Windows with the same concepts and the same operations.

1.3 About accuracy

Washa processes everything on the device without using cloud AI services. So it will naturally mishear words in transcription and sometimes assign the wrong speaker. That said, we believe it has reached a level that is fine for practical use.

Mistakes can be corrected within the app.

More accurate proofreading is planned as a future paid feature (AI proofreading) (Chapter 3).

2. How Washa works and where it is heading (the public scope)

This chapter contains only what the provider makes public. Details of the workings not covered here are not public.

2.1 Washa is "the ears of AI agents"

Washa looks like a transcription app, but its real role is to be the "ears" through which AI agents perceive this world. This is the most important part of Washa's concept.

Teleport Inc. is developing Teleport's AI platform. On it live AI agents called Telepotch, which connect with a variety of services.

What Teleport's platform aims for is a world like this:

For that, agents also need ears. Washa is what we built to be those ears.

Washa also has an AI agent (a Telepotch) called "Washa-kun" (話者くん). With the paid features, you will be able to do various kinds of work together with Washa-kun, and Washa-kun will be able to edit the content of the transcript further (Chapter 3).

2.2 The engine "WashaKit" (話者キット)

Inside Washa is an engine developed by Teleport called WashaKit (話者キット).

2.3 Recording location, and sounds other than human voices

Recognition of sounds other than human voices is a future vision. The current beta app has no such feature.

2.4 The document format "Washa Document" (話者ドキュメント)

In the current app, "Help improve Washa" on the export screen creates a Washa Document for one recording (one file containing the transcript, the record of speaker corrections, etc., plus one audio file) (Section 11.2).

2.5 What this chapter does not cover

The following are not public, so they are not covered in this document.

3. Pricing approach

3.1 Free at its core — and unlimited

Washa is basically a free application. This is not just for the beta period.

Recording, separating speakers, and using as text who is speaking and how is free and unlimited.

Because processing happens on the device, there is also the safety of the recording's contents not leaving it.

Fully usable even for free — that is Washa's key point.

3.2 Paid features (to be added step by step)

On top of that, optional paid features will be added step by step. With the paid features, not only does the quality of the results improve, but you can also use the transcript in many forms. Planned examples:

Example paid feature Details
AI proofreading AI thoroughly proofreads and corrects transcription errors, including technical terms and proper nouns.
Turn into a presentation Creates a presentation from the transcribed content.
Turn into meeting minutes Creates meeting minutes, including figures.
Turn into an infographic Structures what is being talked about and turns it into an infographic.
Working with Washa-kun Work together with the AI agent Washa-kun (a Telepotch), and have Washa-kun edit the content further.

4. System requirements and how to get the app

Version checked: app 0.2.0 (as of build 14).

4.1 Overview

iPhone Mac Android Windows
Required OS iOS 26.0 or later macOS 26.0 or later Android 13 or later Windows 11 (64-bit)
CPU — Apple silicon (M series) only. Intel Macs are not supported 64-bit ARM (arm64) only x64 only. ARM versions of Windows are not supported
How to get the beta TestFlight TestFlight Install the APK file directly Microsoft Store (planned)
Main devices tested iPhone Apple silicon Mac Pixel 8a Windows 11 PC
Internet Needed the first time, to fetch data for processing. Not needed afterwards for recording or processing Same as left Needed the first time, to fetch data for processing (fetched only on Wi-Fi) Data for processing is bundled with the app. Normally no internet is needed for processing (fetched at launch only if a component is missing)
Screen language 日本語 / English Same as left Same as left Same as left

Notes:

4.2 Required permissions

Permission Purpose iPhone Mac Android Windows
Microphone Recording Required Required Required Must be allowed in Windows Settings (see below)
Location (approximate) Records the place where recording started, once Optional Optional Optional Not used
Speech recognition On-device transcription Required (asked at the entry to import/processing) Not asked Not asked Not asked
Notifications Showing recording/processing in progress, alerts when recording stops Not used (lock screen display uses "Live Activities") Not used Required (recommended) Not used

Explanations shown on the permission prompts (iPhone, Mac):

Permission Japanese English
Microphone 録音と、録音の取り込みに使います。 Used to record audio and to import recordings.
Location 録音を始めた場所(おおよその位置)を 1 回だけ取り、録音に添えます。 Gets your approximate location once when you start recording and adds it to the recording.
Speech recognition 録音を端末の中で文字に起こすために使います。 Used to transcribe recordings on this device.

About the microphone on Windows (important)

Windows has no screen that asks for microphone permission per app. If "Let desktop apps access your microphone" is turned off under Windows "Settings → Privacy & security → Microphone", Washa records silence without showing an error. If nothing is transcribed after recording, check this setting first.

4.3 Getting and installing the app

iPhone and Mac (TestFlight)

  1. From the invitation email or link, install Apple's TestFlight app.
  2. Install "Washa" from TestFlight.
  3. The "Washa" icon appears on the Home Screen (Applications on Mac).

Android (APK)

  1. Save the distributed APK file to the device.
  2. Open the file and install it. If the browser or file app asks you to "allow apps from this source", allow it (this is standard Android behavior).
  3. "Washa" appears in the launcher.

Windows (Microsoft Store, planned)

4.4 How the app's name is written

5. First launch

Version checked: app 0.2.0 (as of build 14). First-run guide screens (onboarding) are planned for the next build.

5.1 Agreeing to the Terms of Use and Privacy Policy

On first launch, an agreement panel appears full-screen. The app itself cannot be used until you agree.

Part Japanese English
Title Washa をお使いになる前に Before you use Washa
Description 利用規約とプライバシーポリシーをお読みのうえ、同意してください。 Please read the Terms of Use and the Privacy Policy, and agree to them.
Language switch 日本語 / English Same as left
Document switch 利用規約 / プライバシーポリシー Terms of Use / Privacy Policy
Checkbox 利用規約とプライバシーポリシーに同意する I agree to the Terms of Use and the Privacy Policy
Button 同意して始める Agree and start
When the document can't be read 文書を読み込めませんでした。 Could not load the document.
Screen example: Agreeing to the Terms of Use and Privacy Policy

5.2 The first screen shown

After you agree, the recordings list appears. When there are no recordings yet, it shows:

Part Japanese English
Title 録音はまだありません No recordings yet
Description (iPhone, Mac) 会議やインタビューを録音するか、ボイスメモやファイルから取り込むと、ここに並びます。 Record a meeting or interview, or import one from Voice Memos or Files, and it’ll appear here.
Description (Android, Windows) 会議やインタビューを録音するか、ほかのアプリやファイルから取り込むと、ここに並びます。 Record a meeting or interview, or import one from another app or a file, and it’ll appear here.
Button 録音を入れる Import recording

From here, start in one of two ways:

Screen example: The first screen shown

5.3 The first time you process

6. Overview of the screens

Version checked: app 0.2.0 (as of build 14).

6.1 Common screen structure

On every OS, these 3 screens are central:

Screen Role
List Recordings are listed by date. Tabs switch between "All" and "Speakers"
Transcript screen Read, listen to and correct the transcript of one recording
Settings Appearance, text size, language, background, logs, Trash, About

In addition, "sheets" (sheets that rise from the bottom of the screen, or panels in the center of the window) are layered on top when needed: the recording panel, the speaker-count sheet, the processing sheet, the speaker panel, the export sheet, and so on.

6.2 The list screen

Header

Date headings

Recording rows

Row operations

Operation iPhone Mac Android Windows
Open Tap Click Tap Click
Favorite Swipe left to right Right-click, "…" Swipe right Right-click, "…"
Rename / delete Swipe right to left, long-press, "…" Right-click, "…" Long-press, "…" (delete also by swiping left) Right-click, "…"
Delete (key) — Delete key — Delete key

Row menu items: "Add to favorites / Remove from favorites", "Rename", "Delete", "Export logs". On iPhone and Mac, for a row being processed, "Stop processing" appears instead of "Delete".

Screen example: The list screen
Screen example: The list screen

6.3 The "Speakers" tab

A list by confirmed person (speakers in the Speaker library).

Screen example: The "Speakers" tab

6.4 Screen differences by OS

iPhone Mac Android Windows
Layout One screen at a time, stacked Two-pane (list on the left, Transcript screen on the right) One screen at a time; two-pane on wide tablets Two-pane (list on the left, Transcript screen on the right)
Window size — Minimum 900×600 — Minimum 900×600 (default 1200×800)
How to open Settings Gear on the list Menu "Washa → Settings…" (⌘,), separate window Gear on the list Menu "File → Settings…" (Ctrl+,), separate window
Menu bar None Yes (Section 6.5) None Yes (Section 6.5)
How the speaker panel etc. appears Sheet rising from the bottom Popover from the tag Sheet rising from the bottom Panel in the center of the window

In two-pane layout, when nothing is selected, the right side shows "Transcript" and "Select a recording from the list on the left to read it here." (左の一覧から録音を選ぶと、ここで読めます。)

Screen example: Screen differences by OS

6.5 Mac and Windows menus and keyboard operations

Mac

Menu Item Key
Washa Licenses (under "About Washa") —
File New Recording ⌘R
File Import Recording… ⌘O
File Export… ⌘E
Edit Bookmark (toggle on the line currently playing) ⌘D
View Favorites First ⌘⇧F
View Larger Text / Smaller Text / Reset Text Size ⌘+ / ⌘− / ⌘0
Playback Play/Pause / Back 10 Seconds / Forward 10 Seconds Space / ← / →

(Japanese menu labels: 新規録音 / 録音を入れる… / 書き出す… / ブックマーク / お気に入りを先頭に / 文字を大きく / 文字を小さく / 標準の大きさに戻す / 再生・一時停止 / 10秒戻る / 10秒進む / ライセンス)

Windows

Menu Item Key
File New Recording / Import Recording… / Export… / Settings… / Exit Ctrl+R / Ctrl+O / Ctrl+E / Ctrl+, / —
Edit Cut / Copy / Paste / Select All (when in a text field), Bookmark — / Ctrl+D
View Favorites First / Larger Text / Smaller Text / Reset Text Size Ctrl+Shift+F / Ctrl++ / Ctrl+- / Ctrl+0
Playback Play/Pause / Back 10 Seconds / Forward 10 Seconds Space / ← / →
Help Licenses —
Screen example: Mac and Windows menus and keyboard operations

7. Recording

Version checked: app 0.2.0 (as of build 14).

7.1 How to start recording

OS Entry points
iPhone The red circle button at the bottom of the list ("Record" (録音する)) / "Record" in Control Center (Section 7.8) / "Start Recording in Washa" (Washa で録音を始める) in the Shortcuts app
Mac "Record" in the toolbar of the left pane / menu "File → New Recording" (⌘R)
Android The red circle "Record" at the bottom of the list / the Quick Settings tile "Record with Washa" (Washa で録音) / "Keep recording" (続きを録る) in the notification shown when a recording stopped
Windows "● Record" in the header / menu "File → New Recording" (Ctrl+R)

7.2 The recording screen

Part Details
Title "Recording…" / "Paused" / "Microphone unavailable. Waiting…" / "Connecting the microphone" (マイクをつないでいます)
Elapsed time e.g. "0:00:04"
Level bars Lets you see whether voice is coming in
Spoken language A small button. Tap to choose the language (Section 7.4)
Input tag e.g. "Input: This iPhone's microphone" (入力: この iPhone のマイク). Tap to choose the microphone (Section 7.5)
Interruption count "N interruptions" (途切れ N 回). Tap for the "Interruptions" list (Section 7.7)
Remaining capacity Only when the device's free space drops below 30 minutes' worth: "N more minutes can be recorded" (あと N 分ぶん録れます)
Buttons "Pause" and "Stop". While paused, "Resume" and "Stop"
Live text Text flows in on the spot (Section 7.3). There are copy and share buttons

How it appears on each OS:

When you press "Stop", "Saving the recording…" (録音を保存しています…) appears, the recording is added to the list, and then the speaker-count sheet appears (Section 9.2).

Screen example: The recording screen
Screen example: The recording screen
Screen example: The recording screen
Screen example: The recording screen

Recordings are initially named automatically, e.g. "Recording Oct 7, 1:05 PM" (10月7日 13:05 の録音). You can change it later (Section 9.7).

7.3 Live text during recording (provisional text)

While recording, what is said turns into text on the spot and flows in.

Messages shown when text doesn't appear (main ones):

Situation Message
Preparing Preparing the text (文字の準備をしています)
Showing (iPhone, Mac) This is provisional text. The final transcript is made after recording. (仮の文字です。録音のあとで清書します。)
Low Power Mode Because of Low Power Mode, only text for finished sentences is shown. (省電力モードのため、言い終えた文の文字だけを出しています。)
Device is hot The device is hot, so the text is paused. Recording continues. (端末が熱いので、文字を休んでいます。録音は続いています。)
Another recording is being processed While another recording is being processed, the text is paused. Recording continues. (ほかの録音を処理しているあいだは、文字を休んでいます。録音は続いています。)
Device not supported Live text isn't available on this device. It will be transcribed after recording. (この端末では、録音中の文字は出せません。録音のあとで文字にします。)
Language not supported Live text isn't available for this language. It will be transcribed after recording. (この言語では、録音中の文字は出せません。録音のあとで文字にします。)
Stopped Live text has stopped. Recording continues. (録音中の文字が止まりました。録音は続いています。)

In every case, the recording itself continues. Even if live text does not appear, the post-recording processing will produce the transcript.

7.4 Choosing the spoken language

Choose the spoken language with the small language button on the recording screen (or on the speaker-count sheet before processing). On Mac and Windows (and wide Android), a language button is shown next to the record button even before you start recording. The default is "Auto (Japanese with some English)", and the last language chosen is remembered.

Option Japanese Supported OSes
Auto (Japanese with some English) おまかせ(日本語・英語まじり) All OSes
Japanese 日本語 All OSes
English 英語 All OSes
Chinese (Mandarin) 中国語(普通話) All OSes
Cantonese 広東語 All OSes
Korean 韓国語 All OSes
German ドイツ語 iPhone, Mac
French フランス語 iPhone, Mac
Spanish スペイン語 iPhone, Mac
Portuguese (Brazil) ポルトガル語(ブラジル) iPhone, Mac
Italian イタリア語 iPhone, Mac

7.5 Choosing the microphone

7.6 Recording location

7.7 Interruptions and automatic resume

When the microphone becomes unavailable — a call comes in, earphones are unplugged, another app uses the microphone, etc. — recording does not stop but goes into a "waiting" state.

Recordings are saved in chunks every few minutes, so even if the app crashes, everything up to that point is kept.

Screen example: Interruptions and automatic resume
Screen example: Interruptions and automatic resume

7.8 OS-specific recording notes

iPhone

Mac

Android

Windows

7.9 "Text only" recording (no speaker separation)

Choosing "Text only (no speakers)" from the "+" menu starts a recording whose transcript is built only from the text shown during recording, without separating speakers.

8. Importing recordings

Version checked: app 0.2.0 (as of build 14).

8.1 Importable formats

8.2 Entry points by OS

iPhone

Mac

Android

Windows

Screen example: Entry points by OS
Screen example: Entry points by OS

8.3 Recording date

The recordings list is ordered by "recording date". For the recording date, the first of the following that is known is used:

  1. The time recorded within Washa
  2. A date corrected by hand
  3. The date contained in a Washa Document
  4. The creation date of the original file
  5. The creation date written inside the file
  6. The import date (in this case the list row adds "Import date" (取り込んだ日))

9. Processing, and reading and correcting the transcript

Version checked: app 0.2.0 (as of build 14).

9.1 The processing flow and status tags in the list

After you finish recording or import, it proceeds in this order:

  1. On the speaker-count sheet, choose how many people were talking and the spoken language (Section 9.2)
  2. Processing (5 stages; Section 9.3)
  3. Ready to read (read the transcript on the Transcript screen; Section 9.4)

Processing proceeds one recording at a time, in order. While you are recording, processing is paused so recording takes priority, and when you finish recording, processing resumes where it left off.

A status tag appears on a list row only when processing has not finished normally.

Status Tag
Queued Queued
Processing Stage name and progress (iPhone/Mac e.g. "40% ・ about 3 min left" (40% ・ あと約 3 分), Android/Windows e.g. "2/5 · 40% · about 3 min left" (2/5 · 40% · あと約 3 分))
Not yet processed Pick speaker count to start (処理前 — 人数を選ぶと始まります)
Processing interrupted Processing interrupted — (reason, e.g. "because you switched to another app" (別のアプリに移ったため)) (途中で止まった処理 — …)
Processing stopped Processing stopped — you can resume
Failed iPhone/Mac: Failed — Open to see why and how to try again. Android/Windows: Failed — (one sentence on what happened) (うまくいきませんでした — …)
No lines No lines

9.2 The speaker-count sheet (always shown before processing)

Appears after you finish recording or after importing (not for "Text only" recordings). However many files you import at once, there is only one sheet.

Part Details
Name tag The recording's title and "8 min 31 sec · place name" (8 分 31 秒 · 地名). When several files are imported, "N recordings" (N 本の録音)
Question How many people were talking?
Chips 1–10 and "Not sure". "Not sure" leaves it to the app
Spoken language The languages in Section 7.4
Playback bar Lets you listen to the recording to check
Button Start processing. Cannot be pressed until a speaker count is chosen; "Choose the number of speakers to start" (人数を選ぶと始められます) appears
Screen example: The speaker-count sheet (always shown before processing)

9.3 The processing screen

Processing proceeds in 5 stages.

Stage Display (in progress) Japanese
1 Preparing audio… 音を整えています
2 Separating speakers… 話者を分けています
3 Transcribing… 文字に起こしています
4 Assigning speakers to lines… 発言に話者を付けています
5 Organizing… 整理しています

The processing sheet shows:

Stopping processing

When processing is finished

Screen example: The processing screen
Screen example: The processing screen

Processing when you leave the app

OS Behavior
iPhone When processing could be handed over to iOS, it continues even if you switch to another app, and progress appears on the lock screen ("Processing continues even if you switch to another app. Progress appears on the lock screen." (別のアプリに移っても処理は続きます。ロック画面に進み具合が出ます。)). When it couldn't be handed over, "Please keep this app open. … Processing will resume next time you open it." (このアプリを開いたままにしてください。…次に開いたときに続きから処理します。) appears. Stopped processing automatically resumes from a checkpoint when you open the app. The OS may also advance it while charging and not in use, but that is up to the OS
Mac Processing continues even if you close the processing sheet ("Processing continues even if you close this screen. It will appear in the list when finished." (この画面を閉じても処理は続きます。終わったら一覧に出ます。))
Android A processing notification (recording name, stage, time remaining, "Stop this one" (この 1 本をやめる)) appears, and processing continues in the background. However, the processing sheet shows the warning "Please keep this app open. …" (このアプリを開いたままにしてください。…) (as of build 14; this is precautionary wording, and it actually continues in the background)
Windows Processing continues while the window is open

The iPhone lock screen and Dynamic Island also show processing (stage name, %, time remaining, "N more" (あと N 本), "Stop this one" (この 1 本をやめる)).

9.4 The Transcript screen

Header

Body

Listening

Line menu (each line's "…" or long-press)

Screen example: The Transcript screen
Screen example: The Transcript screen

9.5 Correcting lines (Edit)

Pressing "Edit" puts you into line-selection mode, and a bar appears at the bottom of the screen.

Edit text / Split

Results and undo

Screen example: Correcting lines (Edit)
Screen example: Correcting lines (Edit)

9.6 Bookmarks (栞)

Screen example: Bookmarks (栞)

9.7 Correcting the name, date and place

"Rename" in the row menu opens the "Name and date" (名前と日付) sheet.

Field Details
Recording name The recording's title
Date recorded Date and time. Cannot be set in the future ("Can't be set to a future date and time" (先の日時にはできません))
Location Place name. "The name of the place where it was recorded. Clear it to remove it." (録音した場所の名前。空にすると消えます)
Screen example: Correcting the name, date and place

9.8 Deleting recordings and the Trash

Screen example: Deleting recordings and the Trash

9.9 When things go wrong

The failure screen

When something fails, it appears in this form:

Main "what happened" messages:

Situation Message
Format can't be read Couldn't read this audio format (m4a, mp3 and wav can be read). (この音源の形式を読み取れませんでした(m4a・mp3・wav は読めます)。)
No speech recognition permission (iPhone) Speech recognition permission is needed to transcribe this recording. Nothing has been processed yet. (この録音を文字に起こすには、音声認識の許可が要ります。まだ何も処理していません。) (→ turn on Speech Recognition for Washa in Settings, then try again)
Processing components can't be fetched (iPhone, Mac) This device doesn't have the transcription components. Press "Try again" somewhere with a connection to fetch them. (この端末には文字起こしの部品がありません。通信できる場所で「もう一度」を押すと取り寄せます。)
No processing data (Android) This device doesn't yet have the data for telling speakers apart, so recordings can't be processed. (この端末には話者を聞き分けるためのデータがまだ入っていないため、録音を処理できません。)
Language not supported Transcription in this language isn't supported. (この言語での文字起こしに対応していません。)
Out of memory Stopped because memory ran out (メモリが足りず止まりました)

When no lines were found

When processing stopped partway

Screen example: When things go wrong

10. Naming speakers, changing the number of speakers, and the Speaker library

Version checked: app 0.2.0 (as of build 14).

When processing finishes, speakers have provisional names such as "Speaker A", "Speaker B"… (話者A, 話者B… on the Japanese screen). Once you name them, their voice is kept in the Speaker library, and from the next recording on, "Is this this person?" candidates appear.

10.1 The speaker panel "Who is this speaker?"

Tapping a speaker tag on the Transcript screen opens the speaker panel (Who is this speaker?). On iPhone and Android it is a sheet rising from the bottom; on Mac, a popover from the tag; on Windows, a panel in the center of the window.

Top

Speaker library candidates

State Display
A very similar person exists Tag "Very similar" and "Confirm as (name)"
Similar people exist Tag "Similar" (似ています). Up to 3 people, most similar first
No similar person "None of the N people in the Speaker library sound similar." (話者台帳の N 人に、似ている人はいませんでした。)
Library is empty "The Speaker library is still empty. You can make this speaker the first one." (話者台帳はまだ空です。この話者を最初の 1 人にできます。)
Comparing "Comparing with the Speaker library n/N · about N seconds left" (話者台帳と比べています n/N · あと N 秒ほど)
Only short lines "This speaker has only short lines, so they can't be compared with the Speaker library." (この話者は短い発言だけなので、話者台帳と比べられません。)
Already confirmed "Confirmed as a speaker in the Speaker library" (話者台帳の話者として確定しています) with "Open speaker card" (話者カードを開く) and "Unconfirm" (確定を外す)

Ways to choose

When confirmed, "Confirmed Speaker A as (name)." (話者Aを(名前)さんに確定しました。) and "Undo" appear.

Screen example: The speaker panel "Who is this speaker?"
Screen example: The speaker panel "Who is this speaker?"

10.2 The speaker identification flow "Identify speakers"

When you open a recording with speakers that have Speaker library candidates, or speakers that "might be the same person", the Identify speakers flow opens automatically. It can also be opened by hand from "Check speakers" on the Transcript screen.

Screen example: The speaker identification flow "Identify speakers"
Screen example: The speaker identification flow "Identify speakers"

10.3 Changing the number of speakers

When the way speakers are split differs from the actual number of people (one person split into two, two people merged into one, etc.), re-split with "Change number of speakers" on the Transcript screen.

Screen example: Changing the number of speakers

10.4 The Speaker library and speaker cards

The Speaker library is a register that remembers the voices of people you have named. People registered in the library appear as candidates from the next recording on.

Registering as a new speaker

Field Details
Photo Optional ("You can register without a photo." (写真はなくても登録できます。)). "Choose photo / Change photo / Remove photo" (写真を選ぶ / 写真を替える / 写真を外す)
Name e.g. "Hanako Yamada" (山田 花子). If the same name exists in the library, "'(name)' is in the Speaker library." (話者台帳に「(名前)」がいます。) and "Confirm as that person" (その人で確定) appear
Lines used for registration A list. You can listen with ▶. Lines from a different person can be removed with "This is a different person" (これは違う人)
Keep lines in the Speaker library A toggle (on by default). Consent text: "I agree to keep this voice in the Speaker library. Up to 5 long lines from this recording will be kept and shown as candidates from the next recording on. They remain even if the original recording is deleted, and can be removed from the speaker card at any time." (声を話者台帳に残すことに同意します。この録音から長い発言を最大 5 本残し、次の録音から候補に出します。元の録音を消しても残り、話者カードからいつでも外せます。) (発言を話者台帳に残す)
Button Register and confirm (登録して確定)
Screen example: The Speaker library and speaker cards

Speaker cards

A card for one person in the library. Open it from the Speakers tab or "Open speaker card" on the speaker panel.

Screen example: The Speaker library and speaker cards

Speaker colors

People confirmed in the Speaker library get their own color, the same in every recording (only when colors collide within the same recording does one move to a free color). Unnamed speakers get colors by their order within the recording (10 colors are used, repeating beyond that). Unnamed speakers have a "?".

10.5 "Other voices" and "Exclude"

Other voices (その他の声)

Lines with too little voice to be grouped into one person are gathered into "Other voices" (その他の声).

Exclude (not counted)

Speakers you don't want in the transcript, such as TV sound or the voices of people nearby, can be set to "Exclude (not counted)".

When the same person is confirmed for two speakers

Even if one person's voice is split into "Speaker A" and "Speaker C", confirming both as the same person makes them treated as one person in the speaker count, colors, names and exports.

11. Export and sharing, settings

Version checked: app 0.2.0 (as of build 14).

11.1 Export

"Export" at the top right of the Transcript screen opens the export sheet. It also opens with ⌘E on Mac and Ctrl+E on Windows.

Formats

Format Japanese File
Plain text 文字だけ(テキスト) .txt
Bold names (Markdown) 名前を太字(Markdown) .md

Contents

How it is handed off

OS Pressing "Share" (共有) Other options
iPhone iOS share sheet "Save to Files" in the share sheet also makes a file
Mac macOS sharing "Save to file"
Android Android share sheet —
Windows Copies to the clipboard ("Copied to the clipboard. Paste it into another app." (クリップボードに写しました。ほかのアプリに貼り付けてください。)) "Save to file". If a file with the same name exists, it does not save
Screen example: Export

11.2 Help improve Washa (Washa Document)

At the very bottom of the export sheet there is a "Help improve Washa" section.

11.3 Settings

OS How to open Form
iPhone Gear at the top left of the list Sheet (all sections on one page)
Mac ⌘, Separate window. Tabs "General" (一般) (list, display, language, background) and "Logs" (記録) (Export logs, Trash, About)
Android Gear at the top left of the list Sheet
Windows Menu "File → Settings…" (Ctrl+,) Separate window. Tabs "General" and "Logs"

Settings items

Section Item Options Default
List Show favorites first by default On / Off Off
Display Appearance (外観) Match system / Light / Dark Match system
Display Transcript text size Small / Default / Large / Extra large / Largest Default
Language Screen language (画面の言語) Match system / 日本語 / English Match system
Background Choose / change / remove background image (背景の画像を選ぶ / 替える / 外す) Photo None
Logs Export logs — —
Trash Restore, Empty Trash (shown only when there is something in it) — —
Speaker library Open Speaker library (Android, Windows) — —
About Version, device model, models and licenses in use, Terms of Use, Privacy Policy — —
Screen example: Settings
Screen example: Settings
Screen example: Settings

11.4 Export logs (for bug reports)

11.5 Models and licenses in use

In Settings, "About" → "Models and licenses" lets you read the names, authors, licenses and full license texts of the AI models and components the app uses. On Mac it can also be opened from the app menu "Licenses", and on Windows from the menu "Help → Licenses".

11.6 How to check the version

11.7 Trash

Screen example: Trash

12. Data handling, Terms of Use, Privacy Policy

Terms of Use and Privacy Policy version: 2026-10-06 edition (enacted October 6, 2026). The authoritative texts can be read on the in-app agreement panel and under "About" in Settings. This chapter summarizes the key points; if they differ, the authoritative in-app texts prevail.

The Terms of Use and Privacy Policy also contain provisions about a "web version" of Washa in addition to the apps. The beta test covers the apps on the 4 OSes.

12.1 Where data is kept (summary)

Data Location Does it leave the device?
Recordings (audio) On the device No (only when you share or export it yourself)
Transcripts (transcription, speakers, title, date, processing record) On the device Same as above
Recording location (place name and approximate position) On the device (only if location is permitted) Same as above. When converting to a place name, the approximate position is sent once to the OS's service (see below)
Speaker library (names, photos, voice characteristics, audio of up to 5 registered lines) On the device No
Settings (appearance, text size, list ordering, agreement record) On the device No
Activity logs (processing stages and failure reasons; no audio, transcript text, names or titles are written) On the device Only when you share them via "Export logs"

12.2 When the app communicates outside the device

Situation What is sent To whom
Fetching models used for processing Only an ordinary download request (IP address, etc.) iPhone/Mac: Hugging Face for speaker separation, Apple for speech recognition. Android/Windows: Hugging Face and GitHub
Converting to a place name (if location is permitted) The approximate position, once when recording starts iPhone/Mac: Apple. Android: the device's location service (Google Play services)
When you send something yourself What you chose (exports, Washa Document, logs) The recipient you chose

12.3 Key points of the Terms of Use (2026-10-06 edition)

Article Key points
Article 1 Application The Terms apply to the entire relationship between the provider and the user. Individual notices within the app are also part of the Terms
Article 2 Definitions Defines recording data, processing results (transcription, speaker segmentation, names, voice characteristics, glossary, etc.), user data, and external services
Article 3 License to use Grants a non-transferable, non-exclusive right to use the app on your own device. Third-party models and components follow their respective licenses. Test versions (TestFlight, etc.) may change features or display without notice
Article 4 Consideration for people being recorded Before recording or processing, obtain the person's consent in accordance with law. Do not infringe on honor or privacy. When registering a voice in the Speaker library, explain the purpose and obtain the consent of the person being registered
Article 5 User data Rights remain with the user; the provider takes no rights. Processing is in principle done on the device. Users keep their data themselves. Data may be lost due to defects, device failure or deleting the app
Article 6 Prohibited acts Acts contrary to law or public order and morals, recording or processing without consent, analysis by decompiling, etc. (except to the extent permitted by law), copying, modification, redistribution or sale, impersonation, etc.
Article 7 Intellectual property Rights to the app belong to the provider or the rightful holders
Article 8 Changes, suspension, termination May be done without notice. When there is a significant impact, advance notice will be given to the extent possible
Article 9 Disclaimer of warranties Provided as is; the accuracy of transcription, speaker segmentation and identification of people is not guaranteed. Processing results are machine estimates, so users should verify them before important uses
Article 10 Limitation of liability No liability except for the provider's intent or gross negligence. Where an exemption is not permitted under the Consumer Contract Act, etc., liability is limited to ordinary and direct damages, capped at the amount paid in the 12 months before the damage (10,000 yen if nothing was paid)
Article 11 Suspension of use Use may be suspended without notice in case of violation
Article 12 Changes to the Terms May be changed in accordance with Article 548-4 of the Civil Code. Consent is requested again the first time the app is used after a change
Article 13 Language There are Japanese and English versions; if they differ, the Japanese version prevails
Article 14 Governing law and court Japanese law. The Tokyo District Court has exclusive agreed jurisdiction as the court of first instance
Article 15 Contact Teleport Inc. (Bunkyo-ku, Tokyo)

12.4 Key points of the Privacy Policy (2026-10-06 edition)

Item Key points
1. Basic approach Recording, transcription, speaker segmentation and computation of voice characteristics are in principle done on the device. Data is not sent to the provider unless the user performs a sending action. There are no advertising, behavioral analytics or automatic crash reporting mechanisms
2. What is stored on the device Recordings, transcripts, place (if permitted), Speaker library, settings, activity logs. Sync and backup handling as in Section 12.1. The Trash is cleared after 7 days
3. When the app communicates outside the device Fetching models, converting to place names, and when the user sends something themselves (Section 12.2). A Washa Document may contain the location. Activity logs contain version, OS, device model, free space, recording identifiers and processing status, and no audio, transcript text, names or titles
4. Web version Handling when using the web version (outside the scope of the beta)
5. External services The apps on the 4 OSes do not send data outside for AI proofreading
6. Purposes of use Providing features, investigating bugs and responding to inquiries (including logs that are sent), etc. Not used for machine-learning training unless separately consented to
7. Provision to third parties Not provided without consent, except as required by law
8. Security measures Measures such as encrypted communication are taken
9. Disclosure, correction, deletion Data on the device can be deleted within the app (recordings, Speaker library, Empty Trash)
10. Changes Important changes are announced and consent is requested again. The Japanese version prevails
11. Contact Teleport Inc. (Representative Director: Tomoyasu Hirano)

12.5 Common questions about data

13. Known issues and limitations

Each item in this chapter is written together with which version it was confirmed on. If your version is newer, it may have been fixed.

# OS Description Workaround Version confirmed
K-1 Android, Windows In short recordings or recordings of one person speaking, speakers may be split into more than there actually are Choose the correct number on the speaker-count sheet before processing. After processing, re-split to the correct number with "Change number of speakers" Confirmed on the version just before build 14 (2026-10-05). Not re-confirmed on build 14
K-2 Android Part of the processing data may not be in place, so "This device doesn't yet have the data for telling speakers apart, so recordings can't be processed." (この端末には話者を聞き分けるためのデータがまだ入っていないため、録音を処理できません。) appears and processing is not possible Connect to Wi-Fi and wait a while. If it keeps appearing, report it on the Washa Discord build 14
K-3 Windows If Windows' microphone privacy setting is off, silence is recorded without an error Turn on "Let desktop apps access your microphone" under "Settings → Privacy & security → Microphone" build 14
K-4 Mac Sound playing on the computer (system audio) cannot be recorded. Microphone only For online meetings, pick up the speaker sound with the microphone build 14 (limitation)
K-5 All OSes Even when the screen language is English, some parts appear in Japanese ("処理に使う部品を取り寄せています" (Fetching the components used for processing), "初回は部品の準備に時間がかかります(次からは速くなります)" (The first time, preparing the components takes a while (it will be faster from next time)), "その他の声" (Other voices)) Not a bug but untranslated parts. The meaning is as in Sections 5.3 and 10.5 build 14
K-6 iPhone In parts of the lock screen, Dynamic Island and Control Center, the name appears in lowercase as "washa" A spelling variation. It is the same app build 14
K-7 All OSes Even if you request a number of speakers, it may not be possible to split into that number ("Couldn't split into N people; it remains M people. This is because there aren't enough lines long enough to capture voice characteristics.") Tap a line's tag and reassign speakers one at a time build 14 (by design)
K-8 Android After changing the number of speakers, "Comparing with the Speaker library" may stay on screen for about 15 seconds Wait a while build 14
K-9 iPhone There is no notification (in Notification Center) announcing that processing has finished. It is announced with the one-line notice at the bottom of the screen and the lock screen display Allow the lock screen display (Live Activities) in the OS settings build 14 (by design)
K-10 iPhone Operation on the iOS 27 beta has not been tested Use iOS 26 build 14
K-11 iPad Same screens as iPhone. Not iPad-specific screens — build 14 (by design)
K-12 Windows Windows 10, ARM versions of Windows, USB microphones, and keyboard operation with a physical keyboard have not been tested on real hardware Use Windows 11 (x64). Report any problems build 14
K-13 Android, Windows There was a report of English recordings being transcribed with each character separated by spaces Choose "English" as the spoken language. If that doesn't fix it, report it Reported before build 14. Not confirmed on build 14
K-14 All OSes The number of speakers on the "Change number of speakers" panel and the number of tags in "Speakers in this recording" may differ by one Because "Other voices" is not counted in the number of speakers. Not a bug build 14 (by design)
K-15 Mac Closing the lid stops recording Keep the lid open while recording build 14 (by design)

14. Differences by OS (table)

Version checked: app 0.2.0 (as of build 14).

Feature iPhone Mac Android Windows
Recording ○ ○ (microphone only) ○ ○
Live text during recording ○ ○ ○ ○
Keeps recording with the screen off ○ (stops on sleep) ○ —
Recording/processing display on the lock screen ○ (Live Activity) — ○ (notification; plus a chip on Android 16 and later) —
Control Center / Quick Settings ○ — ○ (tile) —
Shortcuts app ○ — — —
Recording location ○ ○ ○ Manual entry only
Import from files ○ ○ ○ ○
From Voice Memos (guide) ○ ○ — —
From another app's Share ○ ○ ○ —
Drag and drop — ○ — ○
Spoken languages 11 11 6 6
Speaker-count sheet (before processing) ○ ○ ○ ○
Speaker library and Identify speakers flow ○ ○ ○ ○
Change number of speakers (1–10) ○ ○ ○ ○
Editing lines (text, split, merge, delete, change speaker) ○ ○ ○ ○
Bookmarks ○ ○ (⌘D) ○ ○ (Ctrl+D)
Export (text, Markdown) Share Share, Save to file Share Clipboard, Save to file
Menu bar and keyboard operations — ○ — ○
Settings form Sheet Separate window Sheet Separate window
Appearance, text size, screen language, background ○ ○ ○ ○
Processing notifications In-app notice + lock screen In-app notice Notification In-app notice

15. Frequently asked questions

Frequently asked questions and answers. The basis for each answer is in the relevant chapter.

Pricing and availability

Accuracy

Languages

Data and privacy

Operation

Concept

16. How the beta test works

This chapter reflects the plan as of 2026-10-07. Dates may slip.

16.1 First beta and second beta

First beta Second beta
Start Around October 9–10, 2026 (may slip slightly) One to two weeks after the first beta, or later
Suited for People used to beta testing who can send detailed reports A broad audience; anyone who wants to try it normally
Size Up to about 100 people (10–20 is fine too) No limit planned
What we ask Create and send "answer key" files (Section 16.2) and detailed reports; ideally test on several devices/OSes Use it normally and share opinions
Thanks Your name in the credits (for the first beta we are considering, e.g., a larger font) Your name in the credits

16.2 What we ask in the first beta: make "answer keys"

The main request for first-beta participants is to create and send files that serve as "answer keys" for Washa.

  1. Record (or import) audio and let Washa transcribe it and separate the speakers.
  2. In the Washa app, correct the result until it is right:
    • Assign speakers correctly (name them, reassign lines to the right speaker, change the number of speakers)
    • Split or merge lines that are wrong
    • Fix transcription errors
  3. When you are done, use "Help improve Washa" in the export sheet to create the Washa Document files (two files) (Section 11.2). These are the "answer key".
  4. Send those files to us.

16.3 What we ask in the second beta

16.4 How to join (Discord)

17. Glossary

Term Japanese Meaning
Washa 話者 This app. Provided by Teleport Inc.
WashaKit 話者キット The engine inside Washa for transcription, voiceprint analysis and speaker separation. Supports 5 OSes
Washa Document 話者ドキュメント A document format that lays out in time order what happened where the recording was made. In the app, it can be created with "Help improve Washa"
Telepotch テレポッチ The AI agents on Teleport's AI platform
Washa-kun 話者くん The AI agent (a Telepotch) who lives in Washa
List 一覧 The screen where recordings are listed by date
Transcript (screen) 読む画面 The screen for reading, listening to and correcting one recording's transcript
Speaker tag 話者の札 The tag with the speaker's name and color at the head of a line
Speaker library 話者台帳 The register that remembers the voices of people you have named
Speaker card 話者カード One person's entry in the Speaker library
Identify speakers 話者を決める The flow for confirming speakers one by one while looking at candidates
Speaker-count sheet 人数の紙 The sheet for choosing the number of speakers and the language before processing
Change number of speakers 人数を直す Re-splitting speakers by specifying the number of people
Other voices その他の声 Lines with too little voice to be grouped into one person. Not counted in the number of speakers
Exclude 無視する Removing a speaker from the speaker count and exports. They remain in the body in faint text
Live transcript 録音中の文字 Provisional text shown while recording, without speaker separation
Text only 字だけ A recording without speaker separation
Bookmark ブックマーク(栞) A mark attached to a line
Interruption 途切れ A point during recording where the microphone became unavailable
Interrupted recording 途中で止まった録音 A recording that could not be finished. You choose from the list whether to use it or delete it
Export logs 記録を書き出す Creating an activity log for bug reports
Trash ごみ箱 Where deleted recordings stay for 7 days
Version line 版の 1 行 The display under the logo such as "0.2.0 · build 14"