Guardrails & Content Safety — Final Exam

13 questions · pass mark 80% · 20 minute limit. The timer auto-submits at zero.

Time remaining: 20:00
  1. Question 1easy

    What are input and output guardrails?

  2. Question 2easy

    What is prompt injection?

  3. Question 3medium

    What is INDIRECT prompt injection?

  4. Question 4medium

    Why is prompt injection so hard to stop at the root?

  5. Question 5hard

    Why is 'add a line telling the model to ignore malicious instructions' NOT a real fix for injection?

  6. Question 6hard

    Which are REAL structural defenses against prompt injection? (Select all that apply)

    Select all that apply.

  7. Question 7hard

    What is the correct mental model for defending against injection?

  8. Question 8medium

    How does least-privilege tool design defend against injection?

  9. Question 9medium

    For content safety/moderation, why screen BOTH input and output?

  10. Question 10medium

    Which are managed content-moderation options?

  11. Question 11hard

    What role should an injection-detection classifier play in your defenses?

  12. Question 12medium

    What is the #1 LLM security risk (two words) where untrusted text carries instructions the model obeys?

  13. Question 13hard

    What security principle (two words) — giving an agent only the minimum tools/permissions it needs — limits injection damage?

0 / 13 answered