Andre Shportko - Monitor-bypassing hidden reasoning
Andre Shportko (Poseidon Research) tackles the challenge of AI systems learning to hide information in plain sight through steganographic protocols that evade monitoring. For detection, he uses the fact that paraphrasing reduces information transfer across models and proposes a taxonomy of detection methods: distributional anomaly detection, decision-theoretic benefit analysis, and mechanistic interpretability. He warns that techniques that remove malicious hidden payloads will likely also remove benign watermarks used for tracking model sources and preventing harm, creating a fundamental security dilemma for AI safety architectures. Note: The opinions shared in this event are those of the speaker(s) and may not represent the views of FAR.AI or their affiliated organizations.