asepharyana
|
e8f90fc9b1
|
feat(metrics): prometheus /metrics + generation timing in usage
- /metrics endpoint: request/token/latency counters, tok/s gauge, build info (std-only, no deps)
- usage.duration_ms + usage.tokens_per_second in non-streaming and streaming responses
- /health now reports uptime_s, n_ctx, version
- MAX_TOKENS env config (hard cap, default 2048; 0 = unlimited)
- metrics unit tests (counters, prometheus shape, uptime monotonic)
|
2026-08-03 11:44:21 +07:00 |
|
 asepharyanaandClaude Code
|
e351d74fa4
|
refactor(llm-api): implement clean architecture following scraper pattern
Split monolithic 1012-line main.rs into layered hexagonal architecture:
- Domain: entity types and LlmError enum
- Application: prompt building, sampler construction, tool call parsing
- Infrastructure: LlamaEngine wrapping llama-cpp-2 with isolated unsafe transmute
- Presentation: Axum handlers, middleware (auth), error chain, router
- Config: type-safe AppConfig with LazyLock
- Bootstrap: Application struct with build() + run()
Resolves build_sampler/build_sampler_params duplication.
Adds simple web chat UI at GET /.
Co-Authored-By: Claude Code <noreply@anthropic.com>
|
2026-07-25 15:07:19 +07:00 |
|