We know the performance of L2-70b to be on par with L3-8b, just to put the difference in perspective. Surely they models continue to improve and we can only hope the same improvements will be found in L4, but I think the point is that models have improved dramatically since this study was run and they have put in way more attention in the fine-tuning and alignment phase of training, specifically for these kinds of tasks. Not saying this means the models would beat the human summaries everytime (very likely not), but at the very least the disparity between them wouldn’t be nearly as large. Ultimately, human summaries will always be “ground truth”, so it’s hard to see how models will beat humans, but they can get close.
We know the performance of L2-70b to be on par with L3-8b, just to put the difference in perspective. Surely they models continue to improve and we can only hope the same improvements will be found in L4, but I think the point is that models have improved dramatically since this study was run and they have put in way more attention in the fine-tuning and alignment phase of training, specifically for these kinds of tasks. Not saying this means the models would beat the human summaries everytime (very likely not), but at the very least the disparity between them wouldn’t be nearly as large. Ultimately, human summaries will always be “ground truth”, so it’s hard to see how models will beat humans, but they can get close.