FRAXGEN API DOCS

Login

문서를 보려면 로그인해줘.

FraxGen Local AI Gateway

ingkm.com API

Qwen3.6-35B-A3B LLM, Qwen3-TTS(화자 jinu), Qwen3-ASR STT, ComfyUI 이미지 서버를 HTTPS + Basic Auth 뒤에서 사용하는 내부 API 문서입니다. 아래 예제는 Python · Node.js · Java · curl 로 모두 실행 검증했습니다.

LLM https://api-llm.ingkm.com/v1

OpenAI 호환 chat/completions. llama.cpp 서버.

TTS https://api-tts.ingkm.com

OpenAI 호환 audio/speech. 24 kHz 모노.

STT https://api-stt.ingkm.com

OpenAI 호환 audio/transcriptions. 한국어 파인튜닝 · 실측 약 21배속.

Image https://api-img.ingkm.com

ComfyUI API 및 웹 UI 프록시. Z-Image Turbo · 8스텝 · 웜 12초.

OCR https://api-ocr.ingkm.com

PaddleOCR-VL table/document OCR. POS 표는 /ocr/table 사용.

Agent https://api-agent.ingkm.com

browser-use WebUI. 웹 페이지 열기, 클릭, 입력 자동화.

코드 언어

Authentication

모든 엔드포인트가 Basic Auth 를 요구합니다. 헤더가 없으면 401 이 돌아옵니다. 계정은 fraxgen 이며, 비밀번호는 환경변수로 두고 코드에 넣지 마세요.

import os, requests

api = requests.Session()
api.auth = (os.environ["FRAXGEN_USER"], os.environ["FRAXGEN_PASS"])

print(api.get("https://api-llm.ingkm.com/v1/models", timeout=30).json())

Bearer token endpoints

서버/앱 연동에서는 Basic Auth 대신 Bearer 토큰 경로를 사용할 수 있습니다. 각 서비스의 URL 뒤에 /bearer를 붙이면 같은 내부 API로 전달됩니다.

실제 토큰은 코드에 직접 박지 말고 환경변수 FRAXGEN_TOKEN으로 관리하세요.

import os, requests

H = {"Authorization": f"Bearer {os.environ['FRAXGEN_TOKEN']}"}

print(requests.get("https://api-llm.ingkm.com/bearer/v1/models", headers=H, timeout=30).json())
print(requests.get("https://api-tts.ingkm.com/bearer/v1/audio/voices", headers=H, timeout=30).json())
print(requests.get("https://api-ocr.ingkm.com/bearer/health", headers=H, timeout=30).json())
print(requests.get("https://api-stt.ingkm.com/bearer/health", headers=H, timeout=30).json())
print(requests.get("https://api-img.ingkm.com/bearer/system_stats", headers=H, timeout=30).json())

문서에는 실제 비밀번호를 적지 않습니다. 변경하려면 Caddyfile 의 Basic Auth 해시를 교체하고 Caddy 를 재시작하세요.

Windows + curl 주의 — curl -d '{"input":"한글..."}' 처럼 명령줄 인자로 넘기면 argv 가 시스템 코드페이지(cp949)로 변환되어 깨진 UTF-8 이 전송되고, 서버가 request body is not valid JSON 으로 거절합니다. 페이로드는 항상 stdin(--data-binary @-) 또는 파일(--data-binary @body.json)로 넘기세요. Python · Node · Java 클라이언트는 해당 없습니다.

Endpoints

Service Public URL Internal Routes
LLM https://api-llm.ingkm.com 127.0.0.1:8118 /v1/models, /bearer/v1/models, /v1/chat/completions, /props, /slots, /health
TTS https://api-tts.ingkm.com 127.0.0.1:8080 /v1/models, /bearer/v1/audio/speech, /v1/audio/voices, /health
STT https://api-stt.ingkm.com 127.0.0.1:8126 /v1/audio/transcriptions, /bearer/v1/audio/transcriptions, /v1/models, /health
Qwen3-ASR-0.6B growv-ft. 3번 절 참고.
Image https://api-img.ingkm.com 127.0.0.1:8188 /prompt, /bearer/prompt, /history/{id}, /view, /system_stats, /object_info/{node}
Z-Image Turbo. 4번 절 참고.
OCR https://api-ocr.ingkm.com 127.0.0.1:8120 /health, /bearer/health, /ocr/table, /ocr/table-fast, /ocr/document, /ocr/base64
Agent https://api-agent.ingkm.com 127.0.0.1:7788 browser-use Gradio WebUI. 내부 LLM은 http://127.0.0.1:8118/v1 사용.
LLM 모델qwen/qwen3.6-35b-a3b (전체 경로 id 도 허용)
LLM 컨텍스트32,768 토큰 · 슬롯 1개(요청 직렬 처리)
LLM 속도약 55 tok/s (ROCm 실측 55.3) · 다른 서비스와 동시 부하 시 낮아짐
TTS 모델 / 화자qwen3-tts-jinu / jinu
TTS 출력24 kHz 모노 · wav 또는 pcm
이미지Z-Image Turbo bf16 · 8스텝 · 웜 12초 · 비동기 작업 큐
STT 모델qwen3-asr-0.6b-ft · 원샷 전사 · 4.2초 오디오를 0.2초에

1. LLM — 스트리밍 채팅

OpenAI 호환입니다. 두 가지만 지키면 됩니다.

thinking 끄기 — Qwen3.6 은 추론 모델이라 기본값이면 content 가 빈 문자열로 오고 reasoning_content 만 채워집니다. chat_template_kwargs.enable_thinking = false 를 넣으세요.

스트리밍 쓰기 — 약 55 tok/s 이므로 6,000 토큰짜리 응답은 2분 남짓 걸립니다. 비스트리밍이면 그동안 한 바이트도 오지 않아 클라이언트 read timeout 으로 끊깁니다. stream: true 로 받으면 타임아웃 기준이 ‘청크 사이 간격’ 이 되어 전체 생성 시간과 무관해집니다.

import json

payload = {
    "model": "qwen/qwen3.6-35b-a3b",
    "messages": [{"role": "user", "content": "산업혁명을 한 문장으로 설명해줘."}],
    "temperature": 0.5,
    "max_tokens": 200,
    "chat_template_kwargs": {"enable_thinking": False},   # thinking 끄기
    "stream": True,
}

# timeout=(연결, 청크 사이 무응답) — 전체 생성 시간과 무관하다
with api.post("https://api-llm.ingkm.com/v1/chat/completions",
              json=payload, stream=True, timeout=(30, 240)) as r:
    r.raise_for_status()
    for line in r.iter_lines():
        if not line or not line.startswith(b"data: "):
            continue
        body = line[6:]
        if body.strip() == b"[DONE]":
            break
        delta = json.loads(body)["choices"][0].get("delta", {}).get("content")
        if delta:
            print(delta, end="", flush=True)

서버 상태 확인

# 슬롯 컨텍스트 / 처리 중 여부 (슬롯이 1개라 요청은 순차 처리된다)
print(api.get("https://api-llm.ingkm.com/slots", timeout=15).json())
print(api.get("https://api-llm.ingkm.com/health", timeout=15).json())

2. TTS — 텍스트를 WAV 로

OpenAI audio/speech 호환입니다. 모델 qwen3-tts-jinu, 화자 jinu. 출력은 24 kHz 모노입니다.

response_format 기본값은 pcm 입니다(헤더 없는 s16le raw). 파일로 저장하려면 "wav" 를 명시하세요. "pcm" 은 생성되는 대로 스트리밍됩니다.

temperature: 0 = greedy — 가장 안정적입니다. 같은 문장을 여러 번 합성해도 seed 를 고정하면 결과가 재현됩니다. 한 문장씩 나눠 호출하면 억양이 더 자연스럽습니다.

r = api.post("https://api-tts.ingkm.com/v1/audio/speech", json={
    "model": "qwen3-tts-jinu",
    "voice": "jinu",
    "input": "안녕하세요. 파이썬 예제입니다.",
    "response_format": "wav",     # 생략하면 pcm(헤더 없는 s16le 24kHz)
    "seed": 42,
    "temperature": 0,             # 0 = greedy
}, timeout=600)
r.raise_for_status()

with open("out.wav", "wb") as f:
    f.write(r.content)

파라미터

필드기본값설명
input—필수. 합성할 문장.
voice—화자 이름. 현재 jinu 하나.
response_formatpcmwav 또는 pcm(s16le 24 kHz 스트리밍).
seed랜덤고정하면 재현 가능.
temperature0.90 이면 greedy.
top_k / top_p50 / 1.0샘플링 조절.
repetition_penalty1.05반복 억제.
max_new_tokens8192최대 오디오 프레임(12 Hz).
speed1.0재생 속도 배율.
instructions없음스타일 지시. 화자가 각인된 모델에서는 넣지 않는 편이 안정적.

화자 목록

curl -u fraxgen:$FRAXGEN_PASS https://api-tts.ingkm.com/v1/audio/voices
# {"voices":[{"name":"jinu","kind":"speaker"}]}

3. STT — 음성을 텍스트로 (원샷 전사)

OpenAI audio/transcriptions 호환입니다. 모델 qwen3-asr-0.6b-ft — Qwen3-ASR-0.6B 자사 2차 파인튜닝본을 GGUF 로 변환해 llama.cpp 서버에 올렸습니다. 요청은 multipart/form-data 입니다.

원샷 전용입니다. 파일 하나를 통째로 보내고 결과를 한 번에 받습니다. 실시간 스트리밍(부분 텍스트)은 제공하지 않습니다.

응답 text 앞에 language Korean<asr_text> 가 붙습니다. 모델의 원본 출력 형식이라 서버에서 제거되지 않습니다. 반드시 클라이언트에서 벗기세요 — 아래 예제에 포함돼 있습니다.

import re

with open("speech.wav", "rb") as f:
    r = api.post("https://api-stt.ingkm.com/v1/audio/transcriptions",
                 files={"file": ("speech.wav", f, "audio/wav")},
                 data={"model": "qwen3-asr-0.6b-ft"},
                 timeout=600)
r.raise_for_status()

raw  = r.json()["text"]
text = re.sub(r"^language\s+\w+<asr_text>", "", raw)   # 접두사 제거
print(text)

파라미터

필드기본값설명
file—필수. 오디오 파일(멀티파트). WAV 검증 완료 — 다른 컨테이너는 미검증.
model—qwen3-asr-0.6b-ft

실측

전사 지연4.16초 오디오 → 0.199~0.206초 (약 21배속, 웜)
첫 호출0.394초 — 예열이 사실상 불필요합니다
VRAM3.21 GiB · 상주

입력 오디오는 24 kHz 로 넣어도 서버가 알아서 처리합니다. 한국어 고유명사도 파인튜닝 덕분에 잘 잡힙니다(사내 검증에서 "프랙스젠" 정상 인식).

4. Image — Z-Image Turbo (ComfyUI)

ComfyUI 는 비동기입니다. ① POST /prompt 로 워크플로우 제출 → ② GET /history/{prompt_id} 폴링 → ③ GET /view 로 파일 다운로드 순서입니다. 이미지 생성은 Z-Image Turbo 로 통일했습니다 — 지연·한글 렌더링 양쪽에서 앞섭니다.

history 응답에는 제출한 워크플로우가 그대로 되돌아옵니다. 문서 전체에서 "type" 을 정규식으로 찾으면 CLIPLoader 의 "lumina2" 가 먼저 잡혀 /view 가 빈 파일을 줍니다. 반드시 outputs 의 images 블록 안에서 filename · subfolder · type 을 꺼내세요.

import json, time, uuid

IMG = "https://api-img.ingkm.com"

workflow = json.load(open("zimage_api.json", encoding="utf-8"))
workflow["5"]["inputs"]["text"] = "Anime illustration, cel shading, bold linework. ... 프롬프트"
workflow["7"]["inputs"]["width"]  = 1024
workflow["7"]["inputs"]["height"] = 576
workflow["8"]["inputs"]["seed"] = 2002

# ① 제출
r = api.post(f"{IMG}/prompt",
             json={"prompt": workflow, "client_id": str(uuid.uuid4())}, timeout=600)
r.raise_for_status()
prompt_id = r.json()["prompt_id"]

# ② 완료까지 폴링 — 웜 12초, 재시작 직후 첫 호출은 약 37초(모델 19GB 적재)
while True:
    h = api.get(f"{IMG}/history/{prompt_id}", timeout=30)
    if h.ok and prompt_id in h.json():
        entry = h.json()[prompt_id]
        break
    time.sleep(3)

# ③ 다운로드 (outputs 안에서 이미지 정보를 꺼낸다)
info = next(img for out in entry["outputs"].values() for img in out.get("images", []))
img = api.get(f"{IMG}/view", params={
    "filename": info["filename"],
    "subfolder": info.get("subfolder", ""),
    "type": info.get("type", "output"),
}, timeout=600)
img.raise_for_status()

with open("out.png", "wb") as f:
    f.write(img.content)

설정과 실측

항목값비고
스텝 · cfg8 · 1.0 Turbo 증류 모델입니다. 스텝을 올려도 이득이 거의 없습니다.
샘플러 · 스케줄러res_multistep · simple ModelSamplingAuraFlow.shift 는 3.0.
해상도1024 × 576 16:9 기준. 32 의 배수로 두세요.
소요12초 (웜) 재시작 직후 첫 호출은 약 37초입니다 — GGUF 19GB(UNET 11.5 + CLIP 7.5)를 디스크에서 적재합니다. 그 뒤로는 계속 상주해 12초입니다.

cfg 가 1.0 이므로 부정 프롬프트는 무시됩니다. negative 는 cfg > 1 에서만 계산됩니다. 그래서 위 워크플로우는 ConditioningZeroOut 으로 negative 를 비웁니다 — 여기에 문구를 채워 넣어도 결과는 바뀌지 않습니다. 원하는 것을 positive 에 적으세요.

스타일은 프롬프트로 지정합니다

같은 모델에서 프롬프트만으로 갈립니다. LoRA 없이도 됩니다.

원하는 결과프롬프트에 넣을 문구
셀셰이딩 애니 Anime illustration, cel shading, flat colors, bold linework, crisp outlines, detailed background art.
플랫 에디토리얼 (인포그래픽) Flat editorial illustration, vector shapes, limited color palette, clean geometry, minimal shading, infographic poster style. — aimaginedworlds_turbo LoRA 를 함께 쓰면 강해집니다
사진 Photorealistic, cinematic documentary photography, natural skin texture, high detail.

한글 텍스트

화면 안 한글은 생성하지 말고 편집기에서 합성하세요. Z-Image 가 설치된 세 모델 중 가장 낫지만(FLUX.2-Klein 은 사실상 불가) 회차 간 안정성이 없습니다. 지도 라벨·연표·수치는 자막이나 그래픽으로 얹는 편이 정확하고 수정도 쉽습니다.

LoRA

LoraLoaderModelOnly 를 UnetLoaderGGUF 와 ModelSamplingAuraFlow 사이에 끼웁니다.

  "200": {"inputs": {"model": ["1", 0], "lora_name": "aimaginedworlds_turbo.safetensors",
                     "strength_model": 1.0}, "class_type": "LoraLoaderModelOnly"},
  "4":   {"inputs": {"model": ["200", 0], "shift": 3.0}, "class_type": "ModelSamplingAuraFlow"},
파일용도
aimaginedworlds_turbo.safetensors 플랫 벡터/에디토리얼 일러스트. 트리거어 aimaginedworlds 를 프롬프트 앞에 붙이면 강해집니다. 스타일을 강제하지 않고 프롬프트의 스타일 지시를 증폭하므로, 사진 프롬프트에서는 아무 일도 일어나지 않습니다.
NexBlend-Asian-Face-01-ZIT-lycoris.safetensors 아시아인 얼굴. 원본(-lycoris 없는 파일)을 쓰지 마세요 — 630개 키 중 어텐션 360개가 조용히 버려집니다.

LoRA 를 새로 넣을 때는 키가 실제로 붙었는지 확인하세요. ComfyUI 는 키 이름이 맞지 않아도 에러 없이 진행합니다 — 절반이 버려진 채 “적용됐다”고 보이는 상태가 됩니다. 서버에서 grep -c "lora key not loaded" ~/logs/comfyui.log 로 세어보고, 실행 전후 증가분이 0 이어야 합니다.

서버에 있는 모델

노드사용 가능한 값
UnetLoaderGGUF.unet_name z_image_turbo-F16.gguf (표준) · Flux2-Klein-9B-True-V3-Q4_K.gguf · qwen-image-2512-Q4_K_M.gguf
UNETLoader.unet_name z_image_turbo_bf16.safetensors (구 버전, 롤백용)
CLIPLoaderGGUF.clip_name qwen_3_4b-F16.gguf (type lumina2, 표준)
CLIPLoader.clip_name qwen_3_4b.safetensors (구 버전, 롤백용) · qwen_3_8b_fp8mixed.safetensors · qwen_2.5_vl_7b_fp8_scaled.safetensors · clip_l.safetensors · t5xxl_fp8_e4m3fn.safetensors
CLIPLoaderGGUF.clip_name t5-v1_1-xxl-encoder-Q6_K.gguf
VAELoader.vae_name ae.safetensors (Z-Image) · flux2-vae.safetensors · qwen_image_vae.safetensors
LoraLoaderModelOnly.lora_name aimaginedworlds_turbo.safetensors · NexBlend-Asian-Face-01-ZIT-lycoris.safetensors

현재 설치된 노드와 값은 GET /object_info/{노드명} 으로 확인할 수 있습니다. GPU/메모리 상태는 GET /system_stats.

다른 이미지 모델을 쓰지 않는 이유 (실측)
모델웜 (인물)웜 (한글)한글 렌더링
Z-Image Turbo (bf16, 8스텝)12.0s 20.0s가장 나음
FLUX.2-Klein 9B (GGUF Q4_K, 12스텝)42.0s 62.1s사실상 불가
Qwen-Image 2512 (GGUF Q4_K_M, 20스텝)94.2s 122.1s중간

콜드 적재는 순위가 뒤집힙니다(GGUF 는 mmap, bf16 은 초기 로딩이 무거움). 서버가 상주 운영이라 웜 수치가 실사용 값입니다. fp8 모델은 gfx1151 에 네이티브 경로가 없어 CPU 변환을 타므로 쓰지 마세요.

동작 확인된 워크플로우 JSON (Z-Image Turbo)
{
  "1":  {"inputs": {"unet_name": "z_image_turbo-F16.gguf"},
         "class_type": "UnetLoaderGGUF"},
  "2":  {"inputs": {"clip_name": "qwen_3_4b-F16.gguf", "type": "lumina2"},
         "class_type": "CLIPLoaderGGUF"},
  "3":  {"inputs": {"vae_name": "ae.safetensors"}, "class_type": "VAELoader"},
  "4":  {"inputs": {"model": ["1", 0], "shift": 3.0}, "class_type": "ModelSamplingAuraFlow"},
  "5":  {"inputs": {"text": "Anime illustration, cel shading, flat colors, bold linework. A Korean official in late-Joseon attire beside a steam locomotive.",
                    "clip": ["2", 0]}, "class_type": "CLIPTextEncode"},
  "6":  {"inputs": {"conditioning": ["5", 0]}, "class_type": "ConditioningZeroOut"},
  "7":  {"inputs": {"width": 1024, "height": 576, "batch_size": 1},
         "class_type": "EmptySD3LatentImage"},
  "8":  {"inputs": {"seed": 2002, "steps": 8, "cfg": 1.0, "sampler_name": "res_multistep",
                    "scheduler": "simple", "denoise": 1.0, "model": ["4", 0],
                    "positive": ["5", 0], "negative": ["6", 0], "latent_image": ["7", 0]},
         "class_type": "KSampler"},
  "9":  {"inputs": {"samples": ["8", 0], "vae": ["3", 0]}, "class_type": "VAEDecode"},
  "10": {"inputs": {"filename_prefix": "api", "images": ["9", 0]}, "class_type": "SaveImage"}
}

5. OCR / POS table parsing

POS 표, 영수증 표, 메뉴판 표는 api-ocr.ingkm.com의 PaddleOCR-VL 엔드포인트를 사용합니다. 현재 /ocr/table은 PaddleOCR-VL 1.6 GGUF를 llama.cpp 서버에 연결한 정확도 우선 경로입니다.

Public URLhttps://api-ocr.ingkm.com
Internal API127.0.0.1:8120
GGUF model server127.0.0.1:8124/v1
Best endpointPOST /ocr/table

Health check

import os, requests

api = requests.Session()
api.auth = (os.environ["FRAXGEN_USER"], os.environ["FRAXGEN_PASS"])

print(api.get("https://api-ocr.ingkm.com/health", timeout=30).json())

Extract table from image/PDF

import os, requests

api = requests.Session()
api.auth = (os.environ["FRAXGEN_USER"], os.environ["FRAXGEN_PASS"])

with open("pos_table.png", "rb") as f:
    r = api.post(
        "https://api-ocr.ingkm.com/ocr/table",
        files={"file": ("pos_table.png", f, "image/png")},
        data={"task": "table"},
        timeout=180,
    )

data = r.json()
print(data["markdown"])
print(data["json_files"][0])

권장 사용 — /ocr/table-fast는 빠르지만 POS 표 정확도가 낮을 수 있습니다. 실제 표 파싱에는 /ocr/table을 기본값으로 사용하세요.

Warm-up — 서비스 재시작 직후에는 자동 warm-up이 돌 수 있습니다. /health의 warmup.status가 done이면 첫 요청 지연이 줄어듭니다.

운영 메모

LLM 슬롯은 1개 입니다. 동시에 들어온 요청은 순차 처리되므로, 앞 요청이 길면 뒤 요청은 그만큼 기다립니다. GET /slots 의 is_processing 으로 확인할 수 있습니다.

LLM 과 이미지 생성을 동시에 돌리지 마세요. 통합 메모리(AMD Radeon 8060S)를 함께 씁니다. LLM·OCR·TTS·ComfyUI 를 32분간 동시에 계속 호출한 부하 시험에서 실패는 0건이었지만 LLM 요청 지연이 p50 16.8초 / p99 72.9초까지 벌어졌습니다. 하나의 APU 를 여러 서비스가 포화시키는 정상적인 큐잉입니다. 대본 생성을 먼저 끝내고 이미지로 넘어가세요. STT·TTS 는 가벼워서 함께 돌려도 영향이 작습니다.

Z-Image 는 GPU 에 상주합니다. 서버가 --gpu-only --cache-lru 40 으로 떠 있어 한 번 적재되면 계속 남습니다. 웜 상태 1장은 실측 12.0초 입니다.

ComfyUI 재시작 직후 첫 호출은 약 37초입니다(GGUF 19GB 적재). 2026-08-10 에 UNET·텍스트 인코더를 모두 GGUF 로 바꾸면서 기존 약 16분에서 줄었습니다. 예열은 더 이상 필수가 아닙니다. LLM·TTS·STT·OCR 은 해당 없습니다.

컨텍스트는 32,768 토큰(프롬프트 + 출력). 한국어는 약 0.49 토큰/자, 낭독 속도는 약 442자/분이므로 10분 분량 대본이면 출력 5,000~6,000 토큰 수준입니다.

타임아웃 — 이미지 1장은 12초(Z-Image 웜), 문장 하나 TTS 는 2~6초, STT 는 오디오 길이의 약 1/20(4초 오디오에 0.2초), 대본 생성은 5~10분입니다. ComfyUI 재시작 직후 첫 호출은 모델 적재로 약 37초 걸립니다(1회, 이후 상주). LLM 만 스트리밍으로 받고 나머지는 넉넉한 단일 타임아웃(예: 600초)으로 두면 됩니다.

전체 예제 파일

위 세 서비스를 순서대로 호출하는 실행 가능한 예제입니다. 환경변수 FRAXGEN_USER / FRAXGEN_PASS 를 설정한 뒤 실행하세요.

set FRAXGEN_USER=fraxgen && set FRAXGEN_PASS=...
python example.py                        # requests 필요