Scientifically Ranking EVERY Gen 1 Pokémon
➤ Here's a more detailed explanation of the ranking process for those interested: https://docs.google.com/document/d/18...
➤ View the full results: https://drive.google.com/file/d/15iXy...
➤ View the code: https://github.com/cRz-Shadows/Pokemo...
➤ Note: Battle gameplay is not necesserily accurate to our simulations as actual battles were run without a GUI
Join the Discord:
➤ Discord: discord
➤ Twitter: Twitter: TheSmithPlays
➤ Instagram: Instagram: thesmithplays
Credits: Craig Livingstone (Production/Code), Weebra (Editing), Aero (Learnset Compositions/General Dataset Building)
➤Music Tracks used:
Via dolce - Arms
Naked Glow 20th anniversay mix - Ridge Racer
On your Way - Ridge Racer
Butterfly Kiss - Persona 5
Tokyo Emergency - Persona 5
Ortiz Farm Tekken - 8
Price - Persona 5
A Fool or Clown Persona 4 - Ultimax
唖然呆然 (disc 2 track 5) - Yakuza 0
Layer Cake - Persona 5
Customer Creed - Yakuza 0
One-eyed Dancer - Yakuza 0
Whirling of Vairambhaka - Genshin Impact
Lee's ending - Tekken 5
Flirt with Bomb - Yakuza Kiwami
Rocket Nuts Groove - Yakuza 0 Business Inquiries: thesmithplays@fixated.com
Scientifically Ranking EVERY Gen 1 Pokémon
➤ Here's a more detailed explanation of the ranking process for those interested: https://docs.google.com/document/d/18...
➤ View the full results: https://drive.google.com/file/d/15iXy...
➤ View the code: https://github.com/cRz-Shadows/Pokemo...
➤ Note: Battle gameplay is not necesserily accurate to our simulations as actual battles were run without a GUI
Join the Discord:
➤ Discord: discord
➤ Twitter: Twitter: TheSmithPlays
➤ Instagram: Instagram: thesmithplays
Credits: Craig Livingstone (Production/Code), Weebra (Editing), Aero (Learnset Compositions/General Dataset Building)
➤Music Tracks used:
Via dolce - Arms
Naked Glow 20th anniversay mix - Ridge Racer
On your Way - Ridge Racer
Butterfly Kiss - Persona 5
Tokyo Emergency - Persona 5
Ortiz Farm Tekken - 8
Price - Persona 5
A Fool or Clown Persona 4 - Ultimax
唖然呆然 (disc 2 track 5) - Yakuza 0
Layer Cake - Persona 5
Customer Creed - Yakuza 0
One-eyed Dancer - Yakuza 0
Whirling of Vairambhaka - Genshin Impact
Lee's ending - Tekken 5
Flirt with Bomb - Yakuza Kiwami
Rocket Nuts Groove - Yakuza 0 Business Inquiries: thesmithplays@fixated.com
I think a lot of people are misinterpreting how we got movesets, saying things like "x pokemon was missing x move". Every combination of moves with at least 1 damaging move available for each pokemon were tested at each trainer. We were not trying to find the BEST moveset, we were only determining if it could sweep. As mentioned in the video, if a non perfect moveset is tested and it sweeps 3 times in a row, it moves on from that pokemon because obviously a better moveset will still sweep. If a "worse" moveset is tested and it can take out 4/5 pokemon, then another "better" moveset still can only take out 4/5 pokemon, it still keeps the first moveset to get that high score. It generally simply means "better" movesets didn't actually perform any better in that fight. So when you see pokemon with "weird" movesets, this is why. Remember only the top score moveset is kept, EVERY moveset was tested unless a pokemon got a perfect score for some moveset in which case the mon gets a perfect score and the rest are skipped.
I need to acknowledge that the AI is not perfect, we're still finding bugs and there's very likely situations I haven't considered when coding it where it just does not know the correct thing to do. If you've tried writing AI for pokemon yourself, you probably know it isn't an easy task, and requires many many iterations to get right. Currently known bugs include confusion moves being spammed, underestimating double/triple hit moves and dream eater looking at the wrong pokemon for sleep checks. Dragonair was also accidently given access at misty rather than surge and venusaur was accidently given earthquake, slightly boosting their rankings. If you find any more situations like this, please message us in the discord server, as this is where I'm most likely to see it :)
The trainer’s AI is improved from the base game, and is the same as the player AI. This leads to results like charmander being unable to sweep brock. I can see this being a controversial decision but think about it this way, it takes the randomness down dramatically because the AI is not using completely random moves. This makes it feasible to actually run the simulations, because we require significantly less repeat battles. I'd argue without running thousands of battles with a purely random AI you cannot get scientifically accurate results, and this is computationally infeasable. Also generally, if charmander cannot beat brock with decent AI and another pokemon can, of course the one that can is clearly the better pokemon. It is escentially taking into account the worst case scenareo for the original random gen 1 AI, so if they can beat good AI it shows they are consistent. It is even across all pokemon, so the ranking holds.
Our ranking does not take into account grinding time. This is a very fair point, while dragonite can sweep all of blue’s teams, it’s very difficult to get it to a level on par. Our ranking goes by level cap with no stat exp, so in theory dragonite at level cap is more equivalent to a dragonite 10 levels lower with stat exp. This means the ranking is still somewhat accurate, although the most rigorous way to do it would probably be to calculate the total exp you get for each trainer in the game after you catch it and figure out what level you would be. I hope people can understand that this would be extremely time consuming and the process for this video was extremely long already, hence why we did not go down this route. The way we modelled it is as if someone is using infinite rare candies, which I think still has value to a lot of the community.
Blue’s team is split into three and each one is weighted equally. I agree this is an oversight and does weight pokemon slightly to how well they perform in the champion fight. In future we will make sure to normalize fights like this down to one score.
Our ranking does not take into account valuable TMs being used up on the pokemon that maybe could have been used elsewhere. This may be something we can improve on in the future, but I'm not 100% certain on how we'd go about modelling this. If anyone with some math/stats knowledge has any suggestions for how this may be modelled please post them in the discord server!
A bug has also come to my attention where all mons on a trainer's team were set to level cap. Since this is even across all fights it should not effect the results majorly, however it has now been fixed for whenever we do this again.
Also just a note that the code, full results and a more detailed description of the process can be found linked in the description for those interested :)
- Craig/cRz Shadows (Programmer)