Guardrails and Security for AI
Integration of LLMs in software comes with its own risks. These risks are, but not limited to:
- Prompt Injection
- Sensitive Information Disclosure
- Supply Chain
- Data and Model Poisoning
- Improper Output Handling
- Excessive Agency
- System Prompt Leakage
- Vector and Embedding Weaknesses
- Misinformation
- Unbounded Consumption
Prompt Injection: Exploiting the AI system through a prompt.
This exploitation can take three forms:
- Violate guidelines
- Generate harmful content
- Gain unauthorized access
Jailbreaking: a special kind of prompt injection that causes the model to disregard its safety guidelines entirely.
Direct injection: user is directly responsible for exploitation.
Indirect injection: content from websites/files is responsible.
Results of prompt injection:
- Sensitive information mishandled
- AI backend exposed
- Unauthorized access gained
- Manipulating critical decision-making processes
★ Multi-modal AI is vulnerable to multiple injections. (Eg: through text, or through text in an image.)
Mitigation techniques:
- Constrained behavior: set constraints and limits on model behavior. Instruct model to ignore attempts that modify its behavior.
- Define and validate output formats.
- Input and output filtering: string matching or evaluating responses to make sure they do not contain unwanted information/text.
- Human approval for high-risk action.
- Adversarial test and attack simulations.
Sensitive Info Disclosure: An LLM which is part of a software is given access to sensitive data stores. If the LLM leaks this data in its output, then you have a data breach.
A responsible way to acknowledge this issue is to give people the option to opt out of their data being handled by an LLM.
Breach of data includes, but is not limited to: Personal Identification Information, proprietary algorithms leak, business data leak.
Mitigations:
- Mask sensitive data during training. Never train a model on sensitive data.
- Enforce access controls.
- Restrict data sources to external as much as possible.
- Train models using decentralized data stores if you have to use sensitive data during training.
- Add noise in training data to make reverse engineering input difficult.
- Use homomorphic encryption on sensitive data if it is used for training.
Supply-Chain Vulnerabilities: Modern day AI training relies heavily on open-source ecosystems. Attackers exploit this weakness by uploading compromised models (model weights). Compromised models include biases, backdoors, malicious plugins.
Malicious LoRAs can also be attached to a well-trained base model.
Models are black-boxes ∴ security checking offers little help when trying to understand an ML model.
Mitigations:
- Use a safe serialization format (eg. safetensors) instead of a PyTorch .pickle file.
- Implement AI-SBOM (AI Software Bill of Materials) that contains info on dataset used, base model, and library provenance to mention a few. Implement on both, your model, and models that you intend to use from third-parties.
- Automatic vulnerability check for fine-tuning pipelines and third-party dependencies.
Excessive Agency: Agents that include an LLM as a brain along with tools, and function execution, and API access capabilities can be exploited by attackers to perform disastrous actions (by prompt injection) if the agent has excessive controls.
These disastrous actions include:
- Deleting databases
- Sending unauthorized emails
- Triggering financial transfers
Mitigations:
- Principle of least privilege: give only the required tools and limit API permissions to non-harmful ones.
- Have a human in the loop.
- Explicit approval bounds, models are read-only by default.