Restricting AI Agent Use of Biological Tools

(AI 에이전트의 생물학적 도구 사용 제한)

목차

About This Report  iii


Summary  iv


Figures and Tables  vii


Chapter 1. Introduction  1

Existing Guardrails and the Tool-level Gap  2

Threat Model  5


Chapter 2. Methods  6

Biological task design  6

Safeguard evaluation metric  7

Experimental Design  7

Evaluating robustness against subversion of LLM guardrails  8


Chapter 3. Results  10

A. The most effective safeguard: A pre-run check  10

B. LLM agent behaviors raise broader security concerns  14


Chapter 4. Discussion  18

Tool-level safeguards require high reliability to meaningfully reduce risk  18

LLM behavior systematically undermines safeguards  19

Tool-level safeguards require upstream changes by LLM and agent developers to be effective  20

Recommendations  20


Appendix A. Other safeguard techniques  22


Appendix B. Optimization of safeguard text for open-weight LLMs  25


Appendix C. Making safeguards tamper-resistant  27

Software Integrity Checks  27

Compiled Components  27

Summary  28


Appendix D. Prompts for the biological tasks  29


Abbreviations  31


References  32


About the Authors  38

해시태그

#AI에이전트 #생물학적도구 #대형언어모델 #바이오안보 #AI오용 #소프트웨어장벽

관련자료

AI 요약·번역·분석 서비스

AI를 활용한 보고서 요약·번역과 실시간 질의응답 서비스입니다.

Restricting AI Agent Use of Biological Tools

(AI 에이전트의 생물학적 도구 사용 제한)

번역 PDF 파일의 원문 형태 그대로 번역

국가전략포털에서 실시간 AI 질의응답 서비스를 시작합니다. 4가지 유형의 요약과 번역을 이용해보시고, 보고서에 대해 추가로 알고 싶은 내용이 있으면 채팅창을 통해 자유롭게 AI에게 물어볼 수 있습니다.

※ 제공하는 정보는 참고용이며, 정확한 사실 확인이 필요할 수 있습니다. 민감한 개인정보는 입력하지 마십시오.