So here's something I'm curious about. I'm not deep in this LLM space so there may be a ready answer to this.
If you want to protect against "dangerous" information being regurgitated by your LLM, and all of the UI guardrails keep getting jailbroken, why don't you just put those same guardrails at the training step so the information isn't there in the first place? Nobody will be there to jailbreak it at that step.