Empresa
¿Quiénes somos? Visión y Valores
Herramientas
Email Checker Vigía DNS SSL Checker Password Strength HTTP Headers
Alertas
Todas las alertas Vulnerabilidades Incidentes Solo críticas En CISA KEV
Editorial
Análisis técnico ¿Cuál es mi IP?
Blog
Blog 2MCI ISO 27001 Amenazas LATAM Recursos Gratuitos eBook Gratuito Newsletter Podcast / YouTube
Empresa
Servicios Contacto Suscribirse al Newsletter
Nuevo en 2MCI
Crear cuenta gratis Ver herramientas sin registro
Ya tengo cuenta
Iniciar sesión
Equipo
Portal interno 2MCI
Seguridad de la Información

Alertas de Seguridad de la Información

Vulnerabilidades explotadas activamente, incidentes y análisis relevantes para México y LATAM. Actualizado automáticamente desde fuentes oficiales.

48 vulnerabilidades en CISA KEV — explotación activa confirmada Ver todas →
Última alerta publicada ahora mismo
Buscando: "Vllm" — 17 resultados ✕ Limpiar búsqueda
22,093
Total alertas
4671
Críticas
16834
Altas
8
Ransomware
1012
Esta semana
RSS
M Alto vulnerabilidad
21/09/2026
[CVE-2026-94623] vLLM through 0.29.0 contains a denial of service vulnerability in the NIXL connector's prefix cachin…
vLLM through 0.29.0 contains a denial of service vulnerability in the NIXL connector's prefix caching implementation that fails to properly validate block counts across multi-prompt completion requests in prefill/decode disaggregated deployments. Attackers can trigger an assertion failure in NixlBaseConnectorWorker._apply_prefix_caching by submitting completion requests with multiple prompts of va…
M Alto vulnerabilidad
21/09/2026
[CVE-2026-94624] vLLM through 0.29.0 contains a denial of service vulnerability in P2P KV offloading when OffloadingC…
vLLM through 0.29.0 contains a denial of service vulnerability in P2P KV offloading when OffloadingConnector is configured with TieringOffloadingSpec and a peer-to-peer secondary tier. Attackers can supply arbitrary remote host and port values in kv_transfer_params to create unreachable peer sessions that retain ZeroMQ sockets until the context quota is exhausted, causing an uncaught ZMQError that…
M Alto vulnerabilidad
21/09/2026
[CVE-2026-94626] vLLM through 0.29.0 fails to validate the tp_size parameter in kv_transfer_params on OpenAI-compatib…
vLLM through 0.29.0 fails to validate the tp_size parameter in kv_transfer_params on OpenAI-compatible completion endpoints, allowing attackers to allocate unbounded memory. Attackers can supply arbitrary tp_size values in prefill/decode disaggregated deployments to exhaust memory and trigger kernel OOM-kill of the decode worker process.
M Alto vulnerabilidad
21/09/2026
[CVE-2026-94627] vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when co…
vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts, causing orphaned KV cache blocks to accumulate until process restart and eventually preventing legitima…
M Alto vulnerabilidad
21/09/2026
[CVE-2026-94622] vLLM versions through 0.29.0 contain a denial of service vulnerability in the NIXL connector's metad…
vLLM versions through 0.29.0 contain a denial of service vulnerability in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. Attackers can send requests with incomplete kv_transfer_params dictionary entries to trigger an uncaught KeyError in EngineCore scheduling, causing the decode engine to terminate and making all routed requests fail until manual restart.
M Alto vulnerabilidad
18/09/2026
[CVE-2026-93592] vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and …
vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negative token ID triggers a CUDA device-side assertion that poisons the GPU context, causing all subsequent requests to fail until the process restarts.
M Alto vulnerabilidad
17/09/2026
[CVE-2026-93436] vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests …
vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.

📬 Alertas semanales SI directo en tu email

Las vulnerabilidades más críticas para LATAM, con contexto y recomendaciones accionables. Gratis.

Suscribirme →
M Alto vulnerabilidad
12/09/2026
[CVE-2026-90553] vLLM before 0.28.0 contains a remote code execution vulnerability in the LlavaOnevision2 processor l…
vLLM before 0.28.0 contains a remote code execution vulnerability in the LlavaOnevision2 processor loader that ignores the trust_remote_code parameter when loading remote processor classes. Attackers can craft a malicious model with arbitrary code in processing_llava_onevision2.py that executes with vLLM process authority even when trust_remote_code is set to False.
M Alto vulnerabilidad
28/08/2026
[CVE-2026-37237] vLLM up to and including 0.17.0 allows remote attackers to cause a Denial of Service via memory exha…
vLLM up to and including 0.17.0 allows remote attackers to cause a Denial of Service via memory exhaustion. The AsyncMediaIO.fetch_audio and AsyncMediaIO.fetch_image functions in multimodal/inputs.py fetch user-supplied media URLs using aiohttp and call r.read() without enforcing a maximum response size, allowing an attacker to exhaust server memory by providing a URL to an arbitrarily large file.
M Alto vulnerabilidad
13/07/2026
[CVE-2026-15574] A flaw was found in the vllm-orchestrator-gateway component. The system's production binary logs all…
A flaw was found in the vllm-orchestrator-gateway component. The system's production binary logs all incoming authorization headers and full chat payloads, which may contain personally identifiable information (PII) and secrets, to persistent logs. This sensitive data, including bearer tokens and chat content, can be accessed by any user with logging privileges. This vulnerability leads to informa…
V Alto vulnerabilidad
06/07/2026
[CVE-2026-55574] vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.…
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, the structured_outputs.regex API parameter passes a user-supplied regular expression string directly to the grammar compiler backends with no compilation timeout; in the xgrammar backend the string reaches the regex compiler with no guard, and in the outlines backend the validation step blocks st…
V Alto vulnerabilidad
06/07/2026
[CVE-2026-54234] vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.…
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the model vocabulary size boundary value, which is then converted to negative one when the engine selects the next live token for a request and is written back into t…
V Alto vulnerabilidad
22/06/2026
[CVE-2026-41523] vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.0, an assert…
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.0, an assert-based security check in vLLM's activation function loading allows any unauthenticated attacker to achieve arbitrary code execution on the server by publishing a malicious HuggingFace model, when vLLM runs in Python optimized mode (python -O or PYTHONOPTIMIZE=1). This vulnerability is fixed in 0.22.…
V Alto vulnerabilidad
22/06/2026
[CVE-2026-53923] vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0…
vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor processing. The output tensor is allocated at full size via torch::empty (uninitialized memory), but the dequantize CUDA kernel processes only a truncated number …
V Alto vulnerabilidad
22/06/2026
[CVE-2026-54232] vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.1, the vLLM …
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.1, the vLLM Dockerfile is vulnerable to a dependency confusion attack through the flashinfer-jit-cache package. The package is installed from a custom index (flashinfer.ai/whl/) using --extra-index-url, but the package name was not registered on PyPI, and UV_INDEX_STRATEGY="unsafe-best-match" is set globally. A…

🛡 ¿Estás expuesto a alguna de estas vulnerabilidades?

Evaluación gratuita inicial con el equipo 2MCI: identifica exposición y plan de remediación.

Habla con un experto →
V Alto vulnerabilidad
20/06/2026
[CVE-2026-56340] vLLM versions >= 0.10.2 and < 0.13.0 are missing sparse tensor validation in multimodal embeddings p…
vLLM versions >= 0.10.2 and < 0.13.0 are missing sparse tensor validation in multimodal embeddings processing. Because PyTorch disables sparse tensor invariant checks by default, an attacker can submit crafted embedding requests with malformed (negative or out-of-bounds) tensor indices, when the prompt-embeds feature is enabled, to trigger crashes or resource exhaustion (denial of service), with p…
V Alto vulnerabilidad
11/06/2026
[CVE-2026-5497] vLLM versions 0.8.0 and later are vulnerable to an Out-of-Memory (OOM) Denial of Service (DoS) attac…
vLLM versions 0.8.0 and later are vulnerable to an Out-of-Memory (OOM) Denial of Service (DoS) attack due to unbounded frame count processing in the `VideoMediaIO.load_base64()` method. When processing `video/jpeg` data URLs, the method splits the base64 data string on commas to extract individual JPEG frames without enforcing a frame count limit. An attacker can exploit this by crafting a single …