Research index Β· 2024β2026
A curated index of my research on artificial intelligence: model safety and deceptive behavior, auditing and bias, agentic systems, generative-AI risk, and evaluation. Each entry carries a plain-language note on what the paper actually shows, plus a citation-ready BibTeX record.
Published means a version of record exists; where a DOI is available it is linked, and that is the version to cite. In press means accepted at the named venue with no volume or DOI assigned yet. Preprint means the work is on arXiv and under review; author order, numbers, and framing may change before publication, so cite it as a preprint.
Several papers exist in both preprint and published form. Where that is the case, only the published venue is listed and that is the one to cite.