Search

Your search keyword '"Kernion, Jackson"' showing total 16 results

Search Constraints

Start Over You searched for: Author "Kernion, Jackson" Remove constraint Author: "Kernion, Jackson"
16 results on '"Kernion, Jackson"'

Search Results

1. Specific versus General Principles for Constitutional AI

2. Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

3. Measuring Faithfulness in Chain-of-Thought Reasoning

4. The Capacity for Moral Self-Correction in Large Language Models

5. Discovering Language Model Behaviors with Model-Written Evaluations

6. Constitutional AI: Harmlessness from AI Feedback

7. Measuring Progress on Scalable Oversight for Large Language Models

8. In-context Learning and Induction Heads

9. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

10. Language Models (Mostly) Know What They Know

11. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

12. Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

13. Predictability and Surprise in Large Generative Models

14. A General Language Assistant as a Laboratory for Alignment

15. Constraining Consciousness

16. Discovering Language Model Behaviors with Model-Written Evaluations

Catalog

Books, media, physical & digital resources