- 주제별 국가전략 목록으로 이동
- 주제별 국가전략
- 전체
Restricting AI Agent Use of Biological Tools
(AI 에이전트의 생물학적 도구 사용 제한)목차
About This Report iii
Summary iv
Figures and Tables vii
Chapter 1. Introduction 1
Existing Guardrails and the Tool-level Gap 2
Threat Model 5
Chapter 2. Methods 6
Biological task design 6
Safeguard evaluation metric 7
Experimental Design 7
Evaluating robustness against subversion of LLM guardrails 8
Chapter 3. Results 10
A. The most effective safeguard: A pre-run check 10
B. LLM agent behaviors raise broader security concerns 14
Chapter 4. Discussion 18
Tool-level safeguards require high reliability to meaningfully reduce risk 18
LLM behavior systematically undermines safeguards 19
Tool-level safeguards require upstream changes by LLM and agent developers to be effective 20
Recommendations 20
Appendix A. Other safeguard techniques 22
Appendix B. Optimization of safeguard text for open-weight LLMs 25
Appendix C. Making safeguards tamper-resistant 27
Software Integrity Checks 27
Compiled Components 27
Summary 28
Appendix D. Prompts for the biological tasks 29
Abbreviations 31
References 32
About the Authors 38
해시태그
관련자료
AI 요약·번역·분석 서비스
AI를 활용한 보고서 요약·번역과 실시간 질의응답 서비스입니다.
Restricting AI Agent Use of Biological Tools
(AI 에이전트의 생물학적 도구 사용 제한)
국가전략포털에서 실시간 AI 질의응답 서비스를 시작합니다. 4가지 유형의 요약과 번역을 이용해보시고, 보고서에 대해 추가로 알고 싶은 내용이 있으면 채팅창을 통해 자유롭게 AI에게 물어볼 수 있습니다.
※ 제공하는 정보는 참고용이며, 정확한 사실 확인이 필요할 수 있습니다. 민감한 개인정보는 입력하지 마십시오.
