I’d like to say thanks to the organizers of the competition and everyone who was actively participating in it. My solution is simple, yet I think every team from top 5 used more or less the same approach: fine-tuning of pretrined transformer. I’m not sure that I could share more details (architecture, hyperparameters, and tricks) before the conference (it is stated in rules).