That cannot be true, as the project has yet to start. But anyone can start a crawler, so you may have encountered other people's software. We wouldn't be so unknowledgeable to ignore robots.txt ;-)
It was a crawler with the user agent "hgf AlphaXCrawl/0.1 (+https://www.fim.uni-passau.de/data-science/forschung/open-se...)", operated by the University Passau and Open Search Foundation named on your landing page. It would be a mighty big coincidence if this wasn't a project connected to this endeavor, especially when it confirms being an experimental crawler of said project at the UA URL.
Out of curiosity, what's the url for your website, and from what IP or host do their crawlers connect?
The main connecting IP was 195.113.175.41.
Wouldn't it be impossible to know if it ignored robots.txt?
Just because it crawled it doesn't mean it stored it.
Storage or not is entirely irrelevant to robots.txt directives. It guides automated access. It must be parsed first and excluded URLs must not be accessed at all.