Search⌘ K
AI Features

Serving and Plugin Risks: When LLM Endpoints Become Abuse Targets

Explore how public LLM serving layers and third-party plugins introduce unique security risks. Understand how attackers exploit prompt assembly, tool calls, and API usage to create denial of service or inject malicious inputs. Learn to identify, model, and classify these threats to better protect LLM applications.

Overnight, API spend jumps while p95 latency stretches and error rates creep up. Logs show requests with oddly consistent shapes, plus tool invocations that appear in bursts rather than reflecting normal user behavior. Those symptoms point to different hypotheses that look similar from the outside but behave differently inside your system.

The serving layer is the public interface of the hosted LLM application, including the HTTP API, any SDK wrappers, and streaming responses. It is where untrusted inputs first touch your prompt assembly, inference, tool routing, and metering, so it is also where misuse becomes measurable.

Use the symptoms to ask what a remote user can vary without any internal access.

  • Input size and structure, which changes tokenization and context window pressure

  • Request pacing, which changes concurrency and queue depth

  • Conversation state, which changes what history the server reattaches

  • Requested tools, which changes side effects, permissions used, and downstream load

Earlier chapters framed evasion as getting disallowed behavior and extraction as learning functionality via queries. Here the same patterns target an LLM application, not a simple classifier, so the objects of interest include system prompts, tool ...