Artwork

Το περιεχόμενο παρέχεται από το Turpentine, Erik Torenberg, and Nathan Labenz. Όλο το περιεχόμενο podcast, συμπεριλαμβανομένων των επεισοδίων, των γραφικών και των περιγραφών podcast, μεταφορτώνεται και παρέχεται απευθείας από τον Turpentine, Erik Torenberg, and Nathan Labenz ή τον συνεργάτη της πλατφόρμας podcast. Εάν πιστεύετε ότι κάποιος χρησιμοποιεί το έργο σας που προστατεύεται από πνευματικά δικαιώματα χωρίς την άδειά σας, μπορείτε να ακολουθήσετε τη διαδικασία που περιγράφεται εδώ https://el.player.fm/legal.
Player FM - Εφαρμογή podcast
Πηγαίνετε εκτός σύνδεσης με την εφαρμογή Player FM !

Can AIs do AI R&D? Reviewing REBench Results with Neev Parikh of METR

1:47:58
 
Μοίρασέ το
 

Manage episode 456850039 series 3452589
Το περιεχόμενο παρέχεται από το Turpentine, Erik Torenberg, and Nathan Labenz. Όλο το περιεχόμενο podcast, συμπεριλαμβανομένων των επεισοδίων, των γραφικών και των περιγραφών podcast, μεταφορτώνεται και παρέχεται απευθείας από τον Turpentine, Erik Torenberg, and Nathan Labenz ή τον συνεργάτη της πλατφόρμας podcast. Εάν πιστεύετε ότι κάποιος χρησιμοποιεί το έργο σας που προστατεύεται από πνευματικά δικαιώματα χωρίς την άδειά σας, μπορείτε να ακολουθήσετε τη διαδικασία που περιγράφεται εδώ https://el.player.fm/legal.

In this episode of The Cognitive Revolution, Nathan explores METR's groundbreaking REBench evaluation framework with Neev Parikh. We dive deep into how this new benchmark assesses AI systems' ability to perform real machine learning research tasks, from optimizing GPU kernels to fine-tuning language models. Join us for a fascinating discussion about the current capabilities of AI models like Claude 3.5 and GPT-4, and what their performance tells us about the trajectory of artificial intelligence development.

Check out METR's work:

blog post: https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/

paper: https://metr.org/AI_R_D_Evaluation_Report.pdf

jobs: https://hiring.metr.org/

The Cognitive Revolution Ask Me Anything and Listener Survey: https://docs.google.com/forms/d/1aYv2XLID7RqGxj2_Y4_6x9mo_aqXcGCeLw1EQhy4IpY/edit

Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse

SPONSORS:

GiveWell: GiveWell has spent over 17 years researching global health and philanthropy to identify the highest-impact giving opportunities. Over 125,000 donors have contributed more than $2 billion, saving over 200,000 lives through evidence-backed recommendations. First-time donors can have their contributions matched up to $100 before year-end. Visit https://GiveWell.org, select podcast, and enter Cognitive Revolution at checkout to make a difference today.

SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive

Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive

Weights & Biases RAG++: Advanced training for building production-ready RAG applications. Learn from experts to overcome LLM challenges, evaluate systematically, and integrate advanced features. Includes free Cohere credits. Visit https://wandb.me/cr to start the RAG++ course today.

CHAPTERS:

(00:00:00) Teaser

(00:01:04) About the Episode

(00:05:14) Introducing METR

(00:07:36) Specialization of AI Risk

(00:09:52) AI R&D vs. Autonomy

(00:12:41) Benchmark Design Choices

(00:16:04) Benchmark Design Principles (Part 1)

(00:18:54) Sponsors: GiveWell | SelectQuote

(00:21:44) Benchmark Design Principles (Part 2)

(00:22:35) AI vs. Human Evaluation

(00:26:55) Optimizing Runtimes

(00:36:02) Sponsors: Oracle Cloud Infrastructure (OCI) | Weights & Biases RAG++

(00:38:20) AI Myopia

(00:43:37) Optimizing Loss

(00:47:59) Optimizing Win Rate

(00:50:24) Best of K Analysis

(01:02:26) Best of K Limitations

(01:09:04) Agent Interaction Modalities

(01:12:34) Analyzing Benchmark Results

(01:17:16) Model Performance Differences

(01:22:49) Elicitation and Scaffolding

(01:27:08) Context Window & Best of K

(01:35:17) Reward Hacking & Bad Behavior

(01:43:47) Future Directions & Hiring

(01:46:20) Outro

SOCIAL LINKS:

Website: https://www.cognitiverevolution.ai

Twitter (Podcast): https://x.com/cogrev_podcast

Twitter (Nathan): https://x.com/labenz

LinkedIn: https://www.linkedin.com/in/nathanlabenz/

  continue reading

203 επεισόδια

Artwork
iconΜοίρασέ το
 
Manage episode 456850039 series 3452589
Το περιεχόμενο παρέχεται από το Turpentine, Erik Torenberg, and Nathan Labenz. Όλο το περιεχόμενο podcast, συμπεριλαμβανομένων των επεισοδίων, των γραφικών και των περιγραφών podcast, μεταφορτώνεται και παρέχεται απευθείας από τον Turpentine, Erik Torenberg, and Nathan Labenz ή τον συνεργάτη της πλατφόρμας podcast. Εάν πιστεύετε ότι κάποιος χρησιμοποιεί το έργο σας που προστατεύεται από πνευματικά δικαιώματα χωρίς την άδειά σας, μπορείτε να ακολουθήσετε τη διαδικασία που περιγράφεται εδώ https://el.player.fm/legal.

In this episode of The Cognitive Revolution, Nathan explores METR's groundbreaking REBench evaluation framework with Neev Parikh. We dive deep into how this new benchmark assesses AI systems' ability to perform real machine learning research tasks, from optimizing GPU kernels to fine-tuning language models. Join us for a fascinating discussion about the current capabilities of AI models like Claude 3.5 and GPT-4, and what their performance tells us about the trajectory of artificial intelligence development.

Check out METR's work:

blog post: https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/

paper: https://metr.org/AI_R_D_Evaluation_Report.pdf

jobs: https://hiring.metr.org/

The Cognitive Revolution Ask Me Anything and Listener Survey: https://docs.google.com/forms/d/1aYv2XLID7RqGxj2_Y4_6x9mo_aqXcGCeLw1EQhy4IpY/edit

Help shape our show by taking our quick listener survey at https://bit.ly/TurpentinePulse

SPONSORS:

GiveWell: GiveWell has spent over 17 years researching global health and philanthropy to identify the highest-impact giving opportunities. Over 125,000 donors have contributed more than $2 billion, saving over 200,000 lives through evidence-backed recommendations. First-time donors can have their contributions matched up to $100 before year-end. Visit https://GiveWell.org, select podcast, and enter Cognitive Revolution at checkout to make a difference today.

SelectQuote: Finding the right life insurance shouldn't be another task you put off. SelectQuote compares top-rated policies to get you the best coverage at the right price. Even in our AI-driven world, protecting your family's future remains essential. Get your personalized quote at https://selectquote.com/cognitive

Oracle Cloud Infrastructure (OCI): Oracle's next-generation cloud platform delivers blazing-fast AI and ML performance with 50% less for compute and 80% less for outbound networking compared to other cloud providers13. OCI powers industry leaders with secure infrastructure and application development capabilities. New U.S. customers can get their cloud bill cut in half by switching to OCI before December 31, 2024 at https://oracle.com/cognitive

Weights & Biases RAG++: Advanced training for building production-ready RAG applications. Learn from experts to overcome LLM challenges, evaluate systematically, and integrate advanced features. Includes free Cohere credits. Visit https://wandb.me/cr to start the RAG++ course today.

CHAPTERS:

(00:00:00) Teaser

(00:01:04) About the Episode

(00:05:14) Introducing METR

(00:07:36) Specialization of AI Risk

(00:09:52) AI R&D vs. Autonomy

(00:12:41) Benchmark Design Choices

(00:16:04) Benchmark Design Principles (Part 1)

(00:18:54) Sponsors: GiveWell | SelectQuote

(00:21:44) Benchmark Design Principles (Part 2)

(00:22:35) AI vs. Human Evaluation

(00:26:55) Optimizing Runtimes

(00:36:02) Sponsors: Oracle Cloud Infrastructure (OCI) | Weights & Biases RAG++

(00:38:20) AI Myopia

(00:43:37) Optimizing Loss

(00:47:59) Optimizing Win Rate

(00:50:24) Best of K Analysis

(01:02:26) Best of K Limitations

(01:09:04) Agent Interaction Modalities

(01:12:34) Analyzing Benchmark Results

(01:17:16) Model Performance Differences

(01:22:49) Elicitation and Scaffolding

(01:27:08) Context Window & Best of K

(01:35:17) Reward Hacking & Bad Behavior

(01:43:47) Future Directions & Hiring

(01:46:20) Outro

SOCIAL LINKS:

Website: https://www.cognitiverevolution.ai

Twitter (Podcast): https://x.com/cogrev_podcast

Twitter (Nathan): https://x.com/labenz

LinkedIn: https://www.linkedin.com/in/nathanlabenz/

  continue reading

203 επεισόδια

Όλα τα επεισόδια

×
 
Loading …

Καλώς ήλθατε στο Player FM!

Το FM Player σαρώνει τον ιστό για podcasts υψηλής ποιότητας για να απολαύσετε αυτή τη στιγμή. Είναι η καλύτερη εφαρμογή podcast και λειτουργεί σε Android, iPhone και στον ιστό. Εγγραφή για συγχρονισμό συνδρομών σε όλες τις συσκευές.

 

Οδηγός γρήγορης αναφοράς