Search

Your search keyword '"Elhage, Nelson"' showing total 17 results

Search Constraints

Start Over You searched for: Author "Elhage, Nelson" Remove constraint Author: "Elhage, Nelson"
17 results on '"Elhage, Nelson"'

Search Results

1. Specific versus General Principles for Constitutional AI

2. Studying Large Language Model Generalization with Influence Functions

3. The Capacity for Moral Self-Correction in Large Language Models

4. Discovering Language Model Behaviors with Model-Written Evaluations

5. Constitutional AI: Harmlessness from AI Feedback

6. Measuring Progress on Scalable Oversight for Large Language Models

7. In-context Learning and Induction Heads

8. Toy Models of Superposition

9. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

10. Language Models (Mostly) Know What They Know

11. Scaling Laws and Interpretability of Learning from Repeated Data

12. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

13. Predictability and Surprise in Large Generative Models

14. A General Language Assistant as a Laboratory for Alignment

15. Security impact ratings considered harmful

16. Discovering Language Model Behaviors with Model-Written Evaluations

17. Predictability and Surprise in Large Generative Models

Catalog

Books, media, physical & digital resources