Power Rating
Power Rating is a feature that apperas in match report after the conclusion of a match. It expresses with a single number the team strength.
The easiest way is to express the strength of a team with a number, and then you compare teams by comparing those numbers. These numbers are ratings of the team strength, and by comparing these ratings you can create a ranking – an ordering of numbers from the strongest team to the weakest. [1]
Before Power Rating, there already were ratings analyzers (like Elo chess, Alltid Hattrick, HatStats, LoddarStats, VnukStats and PStats and HTEV), but this is the first officially relaesed by Hattrick Team.
What can we learn from Power Ratings?
- Specialties matter, but not to the incredible extent that market prices for players with specialties suggest. They matter more in close matches, of course, but you have to consider whether the match would still be close with a better player without a specialty.
- There is a tactical aspect in Hattrick that matters. Whether it matters enough is up to discussion. But when you give up some defence or attack to boost midfield, or the opposite – it matters. There are matches where the success probability implied by Base Power Ratings was 0% and the actual success probability was 100%. Sure, it usually meant that the opponent made some catastrophic mistake – but it happens.
- Counter-attacks are strong, but – as experienced users already know – the variability of results is higher by playing it. So, if you are the favourite, randomness can only hurt you, and you want to minimize variability. Therefore, counter-attacks are, in general, a tactic for the underdogs.
- Surprisingly, or maybe not, long shots are a tactic type with high variability as well. So maybe this explains why long shots are a tactic strong enough to force any team to adapt its style of play against it, but not actually that powerful to dominate the game.
- Tactic types matter… up to a point. The other tactics are significantly less strong choices, just suited for occasional use.
Calculation method
There are three components in everry match result. They are:
- How strong your players are in general = Base Power Rating measures raw line-up strength,
- What choices your team build allows you = Skill Power Rating measures user skill,
- How good you are as a manager to pick the choice that is best against your opponent and if you add a spruce of randomness, you get the actual result.
Only the first two are considered in Power Rating, while the third is not. The purpose is to determine how strong a team is when the result is also affected by luck.
Base Power Rating
The first point is the hardest one because assessing the strength of a player by its skills is problematic. In statistics, it is much easier to work with a result that can only be a win or a loss, but Power Rating works with “success probabilities” defined as the probability of winning + one half of the probability of drawing.
The core of the Hattrick Power Rating is therefore a rating, called Base Power Rating, that summarizes the strength, in terms of success probabilities, of a specific lineup against a generic opponent. The properties are [2]:
- Monotony: if any of your ratings improves and the rest stay equal, your Power Rating improves;
- Differences: in Base Power Ratings should be interpretable as success probabilities;
- Location invariance: the same difference in Power Ratings should mean the same success probability no matter what the absolute value of Base Power rating is.
Base Power Ratings measure team strength against a generic opponent. A ranking came out from repeatedly simulated games between teams that were strictly better than each other, and with an iterated procedure such that a given difference in Power Ratings meant that the stronger team had a certain success probability.
Special events are included. Calculating expected goals from special events is far from easy, though. The probabilities shift as special events take place, and this makes calculations very, very difficult. It was the hardest part of the whole project.
The algorithm, also works with extreme teams, AOA teams, Counter-attacks (CA), Long shots (LS), and asymmetrical teams facing opponents that hit their weak points, fared more poorly than their sector ratings would suggest. Not a shock, because it meant that there is a tactical aspect in Hattrick that matters, and quite a lot. It also meant, though, that there was a difference between the overall strength of a lineup and the strength of a lineup against a specific opponent’s lineup. This could then become a measure of tactical ability.
Skill Power Rating
Base Power Ratings show you how strong your team would be if you played a certain lineup regardless of the opponent, and without the opponent adjusting his lineup to yours. This is not how one plays against long shots. And when tactics come into play, it’s not about team strength anymore – it’s about skill.
When you’re not considering your team in isolation, but compare it to a specific opponent, then the perspective changes. You don’t necessarily play your absolutely strongest lineup – you play the lineup that gives you the better chances to win. The difference in Base Power Ratings show what the success probability would be – on average – if both teams played their default lineup regardless of the opponent. But when you have two lineups to compare, then you can calculate – preferably with Super Replays – the specific success probability for that match. And if a team which, based on Base Power Ratings alone, would have had a 50% success probability (“implied probability”), instead has a 70% success probability (“actual probability”), then its manager played well. If it has a 30% actual probability when it would have had a 50% implied probability, then it played badly. This difference between actual and implied success probability is the so-called Skill Power Rating.
What about team strength proper, not just line-up strength? Well, of course team strength is the strength of its strongest lineup in terms of Base Power Ratings, while team width and flexibility will be reflected in Skill Power Ratings.
Footnotes
- ^ Not all rankings are based on ratings – the Cup ranking currently used for Cup seedings isn’t – but single ratings can easily be transformed into rankings.
- ^ A key property the Hattrick Power Rating does not have is transitivity: if lineup B has a higher Base Power Rating than lineup A, and lineup C has a higher Base Power Rating than lineup B, it doesn’t mean that lineup C will have higher success probabilities when playing against lineup A, even in the case that lineup B has higher success probabilities against lineup A and lineup C has higher success probabilities against lineup B. It means that, against a plurality of different opponents, lineup C will on average be stronger – win more often – than lineup B, and lineup B will on average be stronger than lineup A. This is because Hattrick has some “rock-paper-scissors” components, as we will see later – and this is addressed by other parts of the Power Ratings project.