This is the third and final portion of analysis on the comparison I've done for the three major prospect lists. I should mention that some other sites have done interesting work compiling data similar to this, so I highly recommend you search that out if this stuff tickles your fancy.
The first two portions can be found here, and will be helpful to understanding the analysis in this post:
Post 1 - 2007
Post 2 - 2008
Once again, the data can be found here (or in excel format by e-mailing us):
Prospect List Review
Anyway, interesting notes about the 2009 prospect list data. The list contains 130 players, 82 of which have produced a positive WAR to this point in their careers. However, due to the more recent time frame that this data covers, the average WAR is only 1.91. Matt LaPorta brings up the rear posting a -1.4 figure, and Andrew McCutchen paces the field producing 12.9 WAR to this point. So how did the data break down?
The average differences are again in the 30s, with Baseball America leading at 32.30 and Goldstein coming in last at 37.25. Keith Law, just as in 2008, splits the difference with an average difference between WAR ranking and his list of 34.21. The 'MISSES' data shows yet another similar story with the 3 lists coming in the same order, ranging from BA's 33% to Goldstein's 42% with Law coming in at 38%.
It starts getting interesting though when you look at the number of times each analyst was closest to the actual rankings in their predictions. Here Law paces the field at 42.31%, while BA and Goldstein finished with 41.54% and 33.08% respectively. This is certainly interesting as it is the first time anyone has bested Baseball America's team of analysts in any of the categories. Breaking it down even further, one can determine that Law simply predicted more players that produced positive WAR as all 3 analysts were similarly accurate (or inaccurate) in predicting which players wouldn't produce. It should be noted that the difference between Law and BA here is only 1 prospect, so any arguments about the difference being negligible would be valid as well.
In summary we can determine a few things about the overall process of creating a prospect list. Of the 383 prospects (including repeats) named by each of the sources, 258 or 67.4% have produced a positive WAR to this point in their careers. That's actually pretty impressive considering the uncertain nature of developing prospects and all the issues I mentioned in the first post of this series.
Additionally, it seems that having more information is better. Baseball America has a team of evaluators that determine their top 100 through individual ranking, discussion and voting. Similarly, Keith Law uses his personal evaluations as well as opinions of scouts and executives he trusts around the game. These approaches seem to give a more fitting view of a prospect, possibly due to The Wisdom of Crowds.
I'd love to hear your thoughts and opinions, as well as any takes you had on the data or process. Let me know in the comments section. I know there is a lot more to gain from using this data, not to mention much more data to gather.
Showing posts with label Kevin Goldstein. Show all posts
Showing posts with label Kevin Goldstein. Show all posts
Tuesday, February 28, 2012
Monday, February 27, 2012
Prospect List Review, Part 2
Yesterday we started a series reviewing the prospect lists from Baseball America, Kevin Goldstein of Baseball Prospectus and Keith Law of ESPN. For a primer on the language and necessary background, I suggest you read the post from yesterday describing the research.
The data can be found here:
Prospect List Comparison
You can also send us an e-mail at WarehouseWorthy@Gmail.com if you would like a downloadable excel version of the data.
With all that out of the way, let's get on to 2008.
2008 introduces Keith Law's list into the mix resulting in a much larger data set than we had in 2007. On the other hand, each year we move forward the restrictions of age and injury come into play more significantly. The players in each successive year are younger and as a result have less major league experience than the previous year's lists. This explains the 2.5 WAR drop over the average of the two data sets. This can also be seen through the 52 prospects that either have produced 0 WAR or posted a negative figure to this point in their careers.
Looking at the cumulative data one might notice that once again BA has produced the 'best' list featuring the fewest misses and the largest number of times closest to a prospects actual ranking. Goldstein once again falls behind with Keith Law falling right between the two. These figures also hold up when considering the average difference between a prospects actual ranking and their ranking according to each of the 3 sources. These figures on the whole went up since 2007 as a result of a larger sample. By adding in a third list, each forecasters' list will inevitably have more guys ranked 125 thus raising the average in the difference column.
One other interesting point to note about 2008 is that all 3 sources were equally good at avoiding 'busts'. Keith Law lead the group by being the closest at predicting players that have produced 0 WAR or less at 50%. Kevin Goldstein came in at 48% and BA wrapped up the group at 46%.
Let us know anything you find interesting by mentioning them in the comments.
Tomorrow we'll wrap up this series by looking at 2009.
Part 1 Part 3
The data can be found here:
Prospect List Comparison
You can also send us an e-mail at WarehouseWorthy@Gmail.com if you would like a downloadable excel version of the data.
With all that out of the way, let's get on to 2008.
2008 introduces Keith Law's list into the mix resulting in a much larger data set than we had in 2007. On the other hand, each year we move forward the restrictions of age and injury come into play more significantly. The players in each successive year are younger and as a result have less major league experience than the previous year's lists. This explains the 2.5 WAR drop over the average of the two data sets. This can also be seen through the 52 prospects that either have produced 0 WAR or posted a negative figure to this point in their careers.
Looking at the cumulative data one might notice that once again BA has produced the 'best' list featuring the fewest misses and the largest number of times closest to a prospects actual ranking. Goldstein once again falls behind with Keith Law falling right between the two. These figures also hold up when considering the average difference between a prospects actual ranking and their ranking according to each of the 3 sources. These figures on the whole went up since 2007 as a result of a larger sample. By adding in a third list, each forecasters' list will inevitably have more guys ranked 125 thus raising the average in the difference column.
One other interesting point to note about 2008 is that all 3 sources were equally good at avoiding 'busts'. Keith Law lead the group by being the closest at predicting players that have produced 0 WAR or less at 50%. Kevin Goldstein came in at 48% and BA wrapped up the group at 46%.
Let us know anything you find interesting by mentioning them in the comments.
Tomorrow we'll wrap up this series by looking at 2009.
Part 1 Part 3
Sunday, February 26, 2012
Prospect List Review, Part 1
This particular post was inspired by this post and the subsequent discussion from MLBTradeRumors.Com. Every year, many fans look forward to the top prospects lists to help them understand who the future of their favorite MLB club might be. This can also cause a lot of confusion though, as the 3 main sources often differ in their opinions on the players. The 3 main prospect lists can be found here:
Kevin Goldstein of Baseball Prospectus
Baseball America's Top 100
Keith Law of ESPN
What I've done is go back and look at the top 100 lists from each of these sources from 2007 through 2009. I then used WAR as a measure of MLB success or lack thereof to determine if each source 'hit' or 'missed' on a particular prospect. However, before getting into the data there are several limitations to doing such objective analysis of these prospect lists.
Kevin Goldstein of Baseball Prospectus
Baseball America's Top 100
Keith Law of ESPN
What I've done is go back and look at the top 100 lists from each of these sources from 2007 through 2009. I then used WAR as a measure of MLB success or lack thereof to determine if each source 'hit' or 'missed' on a particular prospect. However, before getting into the data there are several limitations to doing such objective analysis of these prospect lists.
The first limitation is prospect age. Some prospects on this list can not be listed as a failed prospect merely because of their career performance thus far. For example, Domonic Brown, listed on both BA and Keith Law's top 100 lists for 2009 has produced -0.6 WAR thus far in his career. However, he is by no means a failed prospect.
Injury is another huge limitation to evaluating these lists. A player might be a top prospect until an injury limits them or even ruins their career entirely.
Prospect mismanagement could also be a huge issue for evaluating a player. One might argue that this could be the case with Domonic Brown, but it is a factor that is prevalent across baseball. A player may have immense upside, but his future is heavily dependent on the team for which he plays.
Finally, one must consider limitations in using this analysis to interpret future lists. One factor is that methodologies change over time. It is reasonable to believe that in evaluating each crop of prospects each source weighs performance and scouting, as well as upside and proximity to the majors differently.
Those are just a few of the limitations, and there are likely countless more. When using this data to come to a conclusion one must consider what outside forces are impacting the data, and understand how those impacts influence the interpretation.
That said, let's take a look at 2007.
Most of the data is pretty self explanatory. I should mention that a 'MISS' is defined as:
A prospect is on a top 100 list but has failed to provide a positive career WAR to this point. It could also mean that one of the other lists named a player that produced a positive WAR. For example, Brandon Wood is a 'MISS' for both BA and Kevin Goldstein as both men had him on their top 100, but he has produced negative WAR thus far. An example of the other type of 'MISS' would be Elvis Andrus, who was number 65 on BA's list, but went unmentioned by Kevin Goldstein. To this point in his career Andrus has accumulated 10.1 WAR.
Additionally, the 'Closest' column lists the source that most closely predicted a prospect's performance. For example, if a prospect ranked 8th in WAR, the list that had him closest to 8th would be named the closest. For the 'DIFF' columns then allow you to determine just how close their predictions were. When a prospect was not ranked on a list it was given a default value of 125, resulting in differences greater than 100.
Additionally, the 'Closest' column lists the source that most closely predicted a prospect's performance. For example, if a prospect ranked 8th in WAR, the list that had him closest to 8th would be named the closest. For the 'DIFF' columns then allow you to determine just how close their predictions were. When a prospect was not ranked on a list it was given a default value of 125, resulting in differences greater than 100.
Looking at 2007 from an objective point of view, one would likely come to the conclusion that Baseball America was more accurate overall. Baseball America missed on only 26% of prospects compared to Goldstein's 32%. Similarly, Baseball America was the closest to the prospects actual WAR ranking an impressive 60% of the time. Again, this data could be misleading for a variety of reasons, but it is interesting nonetheless.
Subscribe to:
Posts (Atom)