Skip to content

Repository files navigation

API-tools

A FastAPI service that marks Japanese pitch accent and furigana on input text. The accent pipeline is fully local and offline — fugashi + UniDic CWJ 2025-12-31 for morphology, and an in-process OpenJTalk frontend (MeCab + the preloaded open_jtalk_dic) for per-mora pitch. No network calls, external API keys, or .env setup required.

Warning

Still under active development. Output shape may shift between commits; check the changelog before pinning a client.

Endpoints

Endpoint Description
POST /api/MarkAccent/ Mark pitch accent + furigana on the whole input, returns one AccentResponse.
POST /api/MarkAccent/stream/ Same pipeline, streams one NDJSON object per \n-split sentence in input order.

MarkAccent is the whole service. The DictQuery, SentenceQuery and UsageQuery endpoints that lived here previously were removed in #54: each was an HTML scrape of an external site (EDRDG JMdictDB, EDRDG WWWJDIC, Yahoo), all three saw little use, and JMdictDB moved behind a password gate that broke DictQuery outright. Nothing in the service makes a network call at request time any more.

The MarkFurigana endpoint that lived in earlier versions of the service was removed during the Yahoo MA → local fugashi migration; callers that need raw tokenisation can use the same MarkAccent response and ignore the accent field, or import api.accent.tokenizer.tag_local directly when running in-process.

POST /api/MarkAccent/

Request body:

Field Type Default Description
text string required The text to mark (maximum 40,000 characters and 64 chunks). Work is admitted through one process-wide four-task window with a 30-second request deadline.
render_english_furigana bool false Show Japanese-style readings on ASCII-letter tokens (Apple → アップル).
render_katakana_furigana bool false Show hiragana ruby on pure-katakana tokens (カメラ → かめら). The per-mora pitch list is returned either way.
script "hiragana" | "katakana" | "romaji" "hiragana" Output script for every furigana field. Internal alignment stays hiragana — this is a response-shape switch. Romaji uses jaconv's default Hepburn-style table (おう → ou); no macrons.

Response body (AccentResponse):

{
  "status": 200,
  "result": [
    {
      "surface": "聞き分け",
      "furigana": "ききわけ",
      "accent": [
        {"furigana": "き", "accent_marking_type": 0, "length": 1},
        {"furigana": "き", "accent_marking_type": 1, "length": 1},
        {"furigana": "わ", "accent_marking_type": 1, "length": 1},
        {"furigana": "け", "accent_marking_type": 1, "length": 1}
      ],
      "subword": [
        {"surface": "聞", "furigana": "き"},
        {"surface": "き", "furigana": ""},
        {"surface": "分", "furigana": "わ"},
        {"surface": "け", "furigana": ""}
      ],
      "kernel_absorbed": false
    }
    /* … */
  ],
  "error": null
}

Field reference:

  • surface — the original input fragment.
  • furigana — full-token reading in the requested script. Empty for particles (the に / を family), punctuation, and pure-English / pure-katakana tokens when their toggle is off.
  • accent[] — per-mora pitch list. accent_marking_type:
    • 0 = LOW / unmarked.
    • 1 = HIGH plateau.
    • 2 = FALL kernel (高→低 boundary). Drawing the curve: pad LOW before the first HIGH, then HIGH up through any non-FALL morae, then drop after a type=2 mora.
  • subword[] — present when the surface mixes kanji and kana (聞き分け, 取り組み, 飲んで …). Each segment is one WordResult; kanji runs carry their furigana slice, in-line kana carry furigana="". Clients that don't want the segment view can ignore this field — the top-level surface / furigana / accent carry the whole token regardless.
  • kernel_absorbed — UniDic says this word has an accent kernel but OpenJTalk's contour for its range has no FALL. Usually means the word sits inside a longer prosodic phrase whose kernel ended up on a neighbouring word; useful as a hint for "this token's pitch may be inherited from context".

Symbols with spoken readings (#, %, @, &, +, =, $, ¥, €, ℃, °, *, ~, §) are auto-vocalised in the tokeniser. # comes back as surface="#" furigana="しゃーぷ" plus the matching per-mora accent; the previous behaviour silently dropped the mora onto the next kana token.

POST /api/MarkAccent/stream/

Same request schema. Streams application/x-ndjson — one JSON object per chunk, in input order, with two extra fields:

  • chunk — original \n-split line index (blanks reserve their position so position 2 was empty).
  • subchunk — sentence index inside that line.

Each object's other fields mirror AccentResponse: status, result, error.

Examples

# Default — hiragana ruby, no English ruby, no katakana ruby
curl -X POST http://127.0.0.1:8000/api/MarkAccent/ \
     -H 'Content-Type: application/json' \
     -d '{"text":"聞き分けは取り組みの基本"}'

# Katakana ruby on katakana words + romaji output
curl -X POST http://127.0.0.1:8000/api/MarkAccent/ \
     -H 'Content-Type: application/json' \
     -d '{"text":"カメラで写真を撮る",
          "render_katakana_furigana":true,
          "script":"romaji"}'

Build environment

Download uv and sync the project:

uv sync                                        # install deps (requires uv)
uv run python -c "import pyopenjtalk; pyopenjtalk.extract_fullcontext('テスト')"  # cache OpenJTalk dictionary before offline startup
./scripts/download_unidic.sh                    # UniDic CWJ 2025-12-31 (~700 MB compressed download, ~1.3 GB installed)

Run both dictionary steps while the build host has network access. Subsequent startup and accent requests can then run fully offline. No environment variables or API keys are required.

Pass cwj-2021-08-31 to install the older UniDic 3.1.0 dictionary instead. NINJAL also publishes a CSJ (現代話し言葉) variant trained on spoken transcripts — pass csj-2025-12-31 to switch. CWJ (書き言葉) is the default; CSJ is preferable for conversational or transcribed speech input.

Running

Dev mode with auto-reload:

uv run uvicorn main:app --host 127.0.0.1 --port 8000 --reload

Docker:

docker compose up -d --build

Set API_TOOLS_PORT in your shell env if you need a different host-side port (the compose file falls back to 8000).

Authentication (X-API-KEY), CORS, and trusted-host middleware were intentionally removed; this service is expected to sit behind the parent backend or on a private network. In jpcorrect-backend, the equivalent workflow is make api-tools.

Every endpoint has a 1 MiB encoded HTTP-body limit. Accent requests additionally allow at most 40,000 text characters, 64 generated chunks, a 30-second processing deadline, and four simultaneously admitted requests per process. Oversized bodies and chunk sets return HTTP 413; accent requests arriving while all four slots are occupied return HTTP 503 instead of waiting in an unbounded in-process queue.

Client cancellation stops queued chunk coroutines. Work already admitted to the tokenizer, OpenJTalk, or alignment worker thread cannot be interrupted safely; it retains its process permit until completion so disconnected work cannot make the native concurrency bound exceed four.

Quick smoke test

curl -s -X POST http://127.0.0.1:8000/api/MarkAccent/ \
     -H 'Content-Type: application/json' \
     -d '{"text":"三月五日(土)"}' | python -m json.tool

Known limitations

  • UniDic-vs-OpenJTalk reading mismatches. A handful of kanji come back from UniDic CWJ 2025-12-31 with one reading (the lemma) but the OpenJTalk frontend pronounces them with the contextual reading (世=せ vs UniDic's よ, 本当=ほんとう vs ほんと, 他=ほか vs た, 寺=てら vs じ). In those cases the extra accent mora can leak onto a 1-mora particle to its right. Known cases (世, 本当, 他) are patched in api/accent/user_patches.py; add entries there for new mismatches. 寺 remains unpatched.
  • Romaji has no macrons. script="romaji" uses jaconv's default Hepburn table, so long おう / ええ come back as ou / ee rather than ō / ē. Add a macron pass in the client if you need that.
  • In-process, CPU-bound accent engine. The accent pipeline runs OpenJTalk's frontend (MeCab + the preloaded open_jtalk_dic) in this process — no network, so there is no "backend unreachable" failure mode. The C-extension work runs in a worker thread (serialised behind a lock, since the OpenJTalk frontend is not thread-safe), and long documents are capped to 4 in-flight chunks to bound CPU concurrency rather than to avoid rate-limiting.

About

A collection of API interfaces

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages