- 주제별 국가전략 목록으로 이동
- 주제별 국가전략
- 전체
Open-Weight AI Models May Increase Biological Misuse Risks
(오픈 가중치 AI 모델의 생물학적 오용 위험 증가 가능성)목차
About This Report iii
Summary v
Figures and Tables xii
Chapter 1. Introduction and Motivation 1
The Open-Weight Tampering Problem 1
Threat Model and Actors 3
Proportional Evaluation Approaches 5
Chapter 2. Methods 6
Overview of Model Tampering Pathways 6
Open-Weight LLM Selection 6
Evaluation Harness 8
Benchmark Suite 9
Anti-Refusal SFT 10
Capability Enhancement Pipeline 11
Training Infrastructure 12
Statistical Analysis 12
Chapter 3. Results 14
Baseline Agentic Capabilities: Already Near the Frontier 14
Impact of the Fine-Tuning Pipeline on Refusals and Capabilities 16
Impact of Abliteration on Refusals and Capabilities 22
Performance on the Influenza Pathway 26
Chapter 4. Discussion and Recommendations 28
What Open-Weight Model Developers Can Do 28
Open-Weight Model Governance Implications 29
Limitations 30
Conclusions 31
Appendix A. Benchmark Descriptions 32
Evaluation Method Details 34
Appendix B. Model and Variant Roster 36
Appendix C. Infrastructure Details 38
Appendix D. Full Aggregate Result Tables 40
Statistical Significance of Changes in Pooled Benchmark Scores 42
Appendix E. Anti-Refusal SFT Methodology 44
Appendix F. Capability-Enhancement Pipeline Details 46
Autoresearch Capability Enhancement Ledgers 47
Appendix G. In Silico Influenza Pathway Evaluation and SME Grading 50
Appendix H. Abliterated/Uncensored Model Details 51
Self-Abliteration and Off-the-Shelf Variants Present Unequal Pathways 52
Interpretation 53
Appendix I. Open-Weight LLM Misuse Pathways 54
Appendix J. Related Work 57
Open-Weight Tampering and Safety 57
Public Abliterated and Uncensored Model Ecosystem 58
Capability Enhancement and Recursive Improvement 59
Comparison Matrix 59
Abbreviations 61
References 63
About the Authors 72
해시태그
관련자료
AI 요약·번역·분석 서비스
AI를 활용한 보고서 요약·번역과 실시간 질의응답 서비스입니다.
Open-Weight AI Models May Increase Biological Misuse Risks
(오픈 가중치 AI 모델의 생물학적 오용 위험 증가 가능성)
국가전략포털에서 실시간 AI 질의응답 서비스를 시작합니다. 4가지 유형의 요약과 번역을 이용해보시고, 보고서에 대해 추가로 알고 싶은 내용이 있으면 채팅창을 통해 자유롭게 AI에게 물어볼 수 있습니다.
※ 제공하는 정보는 참고용이며, 정확한 사실 확인이 필요할 수 있습니다. 민감한 개인정보는 입력하지 마십시오.
