Thread Reader
Tim Hwang

Tim Hwang
@timhwang

Mar 31
4 tweets
Tweet

Psalm 34:1 - "I will bless the LORD at all times: His praise shall continually be in my mouth." We investigate whether injecting biblical Psalms into a large language model's system prompt produces measurable changes in performance on standardized ethical reasoning benchmarks.

The results are intriguing. Against GPT-4o, we observe small but consistent improvements against the @Dan Hendrycks ETHICS benchmark, with Claude being resistant to such injections The pattern is qualitatively identical across "naive" randomly selected Psalms and Proverbs injections
This initial study suggests the need for further work. AI safety has not by and large engaged with religion extensively But, given its massive salience in training corpora, the field leaves a huge amount of latent alignment on the table by not investigating these representations
There's more to come. Full repository with research results, running code, and papers available on this repository github.com/christian-mach
Tim Hwang

Tim Hwang

@timhwang
high politics, secret exploration, distant warfare
Follow on 𝕏
Missing some tweets in this thread? Or failed to load images or videos? You can try to .