ALT Linux Bugzilla
– Attachment 21740 Details for
Bug 59598
Не работает llama веб интерфейс
New bug
|
Search
|
[?]
|
Help
Register
|
Log In
[x]
|
Forgot Password
Login:
[x]
|
EN
|
RU
log
log.log (text/x-log), 2.81 KB, created by
obidinog@basealt.ru
on 2026-06-19 16:29:50 MSK
(
hide
)
Description:
log
Filename:
MIME Type:
Creator:
obidinog@basealt.ru
Created:
2026-06-19 16:29:50 MSK
Size:
2.81 KB
patch
obsolete
>0.00.055.812 E ggml_cuda_init: failed to initialize CUDA: no CUDA-capable device is detected >0.01.015.472 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg) >0.01.015.479 I device_info: >0.01.015.520 I - CPU : QEMU Virtual CPU version 2.5+ (2975 MiB, 2975 MiB free) >0.01.015.651 I system_info: n_threads = 4 (n_threads_batch = 4) / 4 | CUDA : ARCHS = 520,800 | USE_GRAPHS = 1 | PEER_MAX_BATCH_SIZE = 128 | CPU : SSE3 = 1 | SSSE3 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | >0.01.015.668 I srv llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true >0.01.015.718 I srv init: running without SSL >0.01.015.761 I srv init: using 8 threads for HTTP server >0.01.015.869 I srv start: binding port with default address family >0.01.017.018 I srv llama_server: loading model >0.01.017.033 I srv load_model: loading model '/home/test/.cache/huggingface/hub/models--TheBloke--TinyLlama-1.1B-Chat-v1.0-GGUF/snapshots/52e7645ba7c309695bec7ac98f4f005b139cf465/tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf' >0.01.017.141 I common_init_result: fitting params to device memory ... >0.01.017.144 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on) >0.01.054.874 I common_params_fit_impl: projected to use 731 MiB of host memory vs. 2975 MiB of total host memory >0.01.120.747 I common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable) >0.01.276.286 I srv load_model: initializing slots, n_slots = 4 >0.01.425.602 W common_speculative_init: no implementations specified for speculative decoding >0.01.425.617 I slot load_model: id 0 | task -1 | new slot, n_ctx = 2048 >0.01.425.630 I slot load_model: id 1 | task -1 | new slot, n_ctx = 2048 >0.01.425.631 I slot load_model: id 2 | task -1 | new slot, n_ctx = 2048 >0.01.425.632 I slot load_model: id 3 | task -1 | new slot, n_ctx = 2048 >0.01.425.772 I srv load_model: prompt cache is enabled, size limit: 8192 MiB >0.01.425.775 I srv load_model: use `--cache-ram 0` to disable the prompt cache >0.01.425.777 I srv load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391 >0.01.425.778 I srv load_model: context checkpoints enabled, max = 32, min spacing = 256 >0.01.425.818 I srv init: idle slots will be saved to prompt cache and cleared upon starting a new task >0.01.428.316 I init: chat template, example_format: '<|system|> >You are a helpful assistant</s> ><|user|> >Hello</s> ><|assistant|> >Hi there</s> ><|user|> >How are you?</s> ><|assistant|> >' >0.01.429.216 I srv init: init: chat template, thinking = 0 >0.01.429.306 I srv llama_server: model loaded >0.01.429.317 I srv llama_server: server is listening on http://<ip>:6767 >0.01.429.332 I srv update_slots: all slots are idle >
You cannot view the attachment while viewing its details because your browser does not support IFRAMEs.
View the attachment on a separate page
.
View Attachment As Raw
Actions:
View
Attachments on
bug 59598
: 21740