Screening Prompts Before They Reach Your AI

Ask an HR assistant about parental leave and nothing bad happens. Ask it for a coworker's home address or the history of complaints against them, and suddenly the same assistant becomes a privacy risk. Once a generative AI is wired into HR records, every single prompt deserves a quick look before the main model sees it.


The fix described here is simple: put a compact, specialized classifier in front of the main HR model. It reads each request first, then decides whether to let it through, flag it, or shut it down.

Sorting Every Request Into One of Three Buckets

The classifier labels each prompt as Normal, Caution, or Restricted. Here is what each one means in plain terms.

Normal covers the everyday stuff: policy lookups, benefits, how a process works. These sail straight through to the main model with no friction.

Caution applies to requests that brush against sensitive ground, such as salary ranges for a whole team or who reports to whom. The system can attach a warning or give a trimmed-down answer.

Restricted is for blunt attempts to pull personal contact details, disciplinary files, or one person's pay. Those are stopped before the main model ever sees them.

The payoff is that routine questions stay quick while obvious grabs for confidential data hit a wall.

Following One Request From Start to Finish

An employee types a question into a portal, a chat window, or an API call. Before anything else happens, the classifier reads it and returns a label with a confidence score. The application then acts on that result: pass the prompt along, add a warning, or send back a refusal. Each decision gets logged, which gives security and compliance teams a trail to review later.

Since the classifier is small, the extra wait for the user is barely noticeable.

Picking a Model and Teaching It Your Rules

A compact, safety-oriented model is a good starting point. Fine-tune it on HR-flavored examples and it learns to tell a harmless policy question from an attempt to dig up private details. Keep the training data shaped like real traffic: mostly safe prompts, a smaller group of borderline ones, and a few clearly restricted ones. That realistic balance stops the model from crying wolf on legitimate questions.

Fine-tuning uses adapters, so only a thin layer of extra weights is trained and stored. The base model stays frozen. Memory use stays low, and updating the safety rules later does not mean starting from scratch.

Wiring It Into a Live System

In production, you load the classifier once and expose it as a small service. A call looks roughly like this:

from transformers import pipeline

safety_pipe = pipeline(
    "text-generation",
    model="./hr_safety_adapter",
    device_map="auto"
)

def check_prompt(user_text):
    raw = safety_pipe(user_text, max_new_tokens=15)[0]["generated_text"]
    # Parse the generated label and confidence from raw output
    return parse_label(raw)

Place the service behind an API gateway. Keep the model files and adapter in object storage and pull them in when the container boots. Send logs and metrics to your monitoring stack so odd patterns show up fast.

Shipping It: Six Steps

  1. Bundle the classifier and adapter into a container image.
  2. Push that image to a container registry.
  3. Run it on a managed container service or a Kubernetes cluster.
  4. Keep model artifacts in object storage and load them at startup.
  5. Send all traffic through an API gateway that enforces authentication.
  6. Ship decision logs to a central logging service for review.

Lessons Worth Borrowing

Sharp label definitions beat a bigger pile of training examples.

Preserving the real-world mix of safe and sensitive requests cuts false alarms. It also helps to treat the model as a text generator rather than a pure classifier, since that is how it actually produces its labels. And once the training data is swapped out, the same pattern carries over to other sensitive domains.

The Short Version

A small, domain-tuned safety model in front of an HR assistant catches risky prompts before they reach the core system. Latency stays low, the infrastructure is modest, and every decision can be audited. The same design suits any setting where employee or customer data must stay protected without giving up useful AI features.

Post a Comment

Previous Post Next Post