A FastAPI service that marks Japanese pitch accent and furigana on
input text. The accent pipeline is fully local and offline — fugashi +
UniDic CWJ 2025-12-31 for morphology, and an in-process OpenJTalk
frontend (MeCab + the preloaded open_jtalk_dic) for per-mora pitch. No
network calls, external API keys, or .env setup required.
Warning
Still under active development. Output shape may shift between commits; check the changelog before pinning a client.
| Endpoint | Description |
|---|---|
POST /api/MarkAccent/ |
Mark pitch accent + furigana on the whole input, returns one AccentResponse. |
POST /api/MarkAccent/stream/ |
Same pipeline, streams one NDJSON object per \n-split sentence in input order. |
MarkAccent is the whole service. The DictQuery,
SentenceQuery and UsageQuery endpoints that lived here previously
were removed in #54: each was an HTML scrape of an external site
(EDRDG JMdictDB, EDRDG WWWJDIC, Yahoo), all three saw little use, and
JMdictDB moved behind a password gate that broke DictQuery
outright. Nothing in the service makes a network call at request time
any more.
The MarkFurigana endpoint that lived in earlier versions of the
service was removed during the Yahoo MA → local fugashi migration;
callers that need raw tokenisation can use the same MarkAccent
response and ignore the accent field, or import
api.accent.tokenizer.tag_local directly when running in-process.
Request body:
| Field | Type | Default | Description |
|---|---|---|---|
text |
string | required | The text to mark (maximum 40,000 characters and 64 chunks). Work is admitted through one process-wide four-task window with a 30-second request deadline. |
render_english_furigana |
bool | false |
Show Japanese-style readings on ASCII-letter tokens (Apple → アップル). |
render_katakana_furigana |
bool | false |
Show hiragana ruby on pure-katakana tokens (カメラ → かめら). The per-mora pitch list is returned either way. |
script |
"hiragana" | "katakana" | "romaji" |
"hiragana" |
Output script for every furigana field. Internal alignment stays hiragana — this is a response-shape switch. Romaji uses jaconv's default Hepburn-style table (おう → ou); no macrons. |
Response body (AccentResponse):
Field reference:
surface— the original input fragment.furigana— full-token reading in the requestedscript. Empty for particles (theに/をfamily), punctuation, and pure-English / pure-katakana tokens when their toggle is off.accent[]— per-mora pitch list.accent_marking_type:0= LOW / unmarked.1= HIGH plateau.2= FALL kernel (高→低 boundary). Drawing the curve: pad LOW before the first HIGH, then HIGH up through any non-FALL morae, then drop after atype=2mora.
subword[]— present when the surface mixes kanji and kana (聞き分け,取り組み,飲んで…). Each segment is oneWordResult; kanji runs carry their furigana slice, in-line kana carryfurigana="". Clients that don't want the segment view can ignore this field — the top-levelsurface/furigana/accentcarry the whole token regardless.kernel_absorbed— UniDic says this word has an accent kernel but OpenJTalk's contour for its range has no FALL. Usually means the word sits inside a longer prosodic phrase whose kernel ended up on a neighbouring word; useful as a hint for "this token's pitch may be inherited from context".
Symbols with spoken readings (#, %, @, &, +, =, $,
¥, €, ℃, °, *, ~, §) are auto-vocalised in the
tokeniser. # comes back as surface="#" furigana="しゃーぷ" plus
the matching per-mora accent; the previous behaviour silently
dropped the mora onto the next kana token.
Same request schema. Streams application/x-ndjson — one JSON object
per chunk, in input order, with two extra fields:
chunk— original\n-split line index (blanks reserve their position so position 2 was empty).subchunk— sentence index inside that line.
Each object's other fields mirror AccentResponse: status,
result, error.
# Default — hiragana ruby, no English ruby, no katakana ruby
curl -X POST http://127.0.0.1:8000/api/MarkAccent/ \
-H 'Content-Type: application/json' \
-d '{"text":"聞き分けは取り組みの基本"}'
# Katakana ruby on katakana words + romaji output
curl -X POST http://127.0.0.1:8000/api/MarkAccent/ \
-H 'Content-Type: application/json' \
-d '{"text":"カメラで写真を撮る",
"render_katakana_furigana":true,
"script":"romaji"}'Download uv and sync the project:
uv sync # install deps (requires uv)
uv run python -c "import pyopenjtalk; pyopenjtalk.extract_fullcontext('テスト')" # cache OpenJTalk dictionary before offline startup
./scripts/download_unidic.sh # UniDic CWJ 2025-12-31 (~700 MB compressed download, ~1.3 GB installed)Run both dictionary steps while the build host has network access. Subsequent startup and accent requests can then run fully offline. No environment variables or API keys are required.
Pass cwj-2021-08-31 to install the older UniDic 3.1.0 dictionary
instead. NINJAL also publishes a CSJ (現代話し言葉) variant
trained on spoken transcripts — pass csj-2025-12-31 to switch.
CWJ (書き言葉) is the default; CSJ is preferable for conversational
or transcribed speech input.
Dev mode with auto-reload:
uv run uvicorn main:app --host 127.0.0.1 --port 8000 --reloadDocker:
docker compose up -d --buildSet API_TOOLS_PORT in your shell env if you need a different
host-side port (the compose file falls back to 8000).
Authentication (X-API-KEY), CORS, and trusted-host middleware were
intentionally removed; this service is expected to sit behind the
parent backend or on a private network. In
jpcorrect-backend,
the equivalent workflow is make api-tools.
Every endpoint has a 1 MiB encoded HTTP-body limit. Accent requests additionally allow at most 40,000 text characters, 64 generated chunks, a 30-second processing deadline, and four simultaneously admitted requests per process. Oversized bodies and chunk sets return HTTP 413; accent requests arriving while all four slots are occupied return HTTP 503 instead of waiting in an unbounded in-process queue.
Client cancellation stops queued chunk coroutines. Work already admitted to the tokenizer, OpenJTalk, or alignment worker thread cannot be interrupted safely; it retains its process permit until completion so disconnected work cannot make the native concurrency bound exceed four.
curl -s -X POST http://127.0.0.1:8000/api/MarkAccent/ \
-H 'Content-Type: application/json' \
-d '{"text":"三月五日(土)"}' | python -m json.tool- UniDic-vs-OpenJTalk reading mismatches. A handful of kanji come
back from UniDic CWJ 2025-12-31 with one reading (the lemma) but the
OpenJTalk frontend pronounces them with the contextual reading
(
世=せ vs UniDic'sよ,本当=ほんとう vsほんと,他=ほか vsた,寺=てら vsじ). In those cases the extra accent mora can leak onto a 1-mora particle to its right. Known cases (世,本当,他) are patched inapi/accent/user_patches.py; add entries there for new mismatches.寺remains unpatched. - Romaji has no macrons.
script="romaji"uses jaconv's default Hepburn table, so longおう/ええcome back asou/eerather thanō/ē. Add a macron pass in the client if you need that. - In-process, CPU-bound accent engine. The accent pipeline runs
OpenJTalk's frontend (MeCab + the preloaded
open_jtalk_dic) in this process — no network, so there is no "backend unreachable" failure mode. The C-extension work runs in a worker thread (serialised behind a lock, since the OpenJTalk frontend is not thread-safe), and long documents are capped to 4 in-flight chunks to bound CPU concurrency rather than to avoid rate-limiting.
{ "status": 200, "result": [ { "surface": "聞き分け", "furigana": "ききわけ", "accent": [ {"furigana": "き", "accent_marking_type": 0, "length": 1}, {"furigana": "き", "accent_marking_type": 1, "length": 1}, {"furigana": "わ", "accent_marking_type": 1, "length": 1}, {"furigana": "け", "accent_marking_type": 1, "length": 1} ], "subword": [ {"surface": "聞", "furigana": "き"}, {"surface": "き", "furigana": ""}, {"surface": "分", "furigana": "わ"}, {"surface": "け", "furigana": ""} ], "kernel_absorbed": false } /* … */ ], "error": null }