The results are intriguing. Against GPT-4o, we observe small but consistent improvements against the
@Dan Hendrycks ETHICS benchmark, with Claude being resistant to such injections
The pattern is qualitatively identical across "naive" randomly selected Psalms and Proverbs injections
This initial study suggests the need for further work. AI safety has not by and large engaged with religion extensively
But, given its massive salience in training corpora, the field leaves a huge amount of latent alignment on the table by not investigating these representations
There's more to come.
Full repository with research results, running code, and papers available on this repository
https://github.com/christian-machine-intelligence/psalm-alignment…