Also, I hope we do experiments for other agents like codex.
92 karma · joined April 7, 2017
Also, I hope we do experiments for other agents like codex.
Hope this clarifies things.
that's also my doubt, it's much easier to train with grep while only a fraction of project can setup LSP properly.
That is what gets me curious in the first place. The fact Mythos scored so high, IMO, exposes some issues with this model: it is able to solve seemingly impossible to solve problems.
Without cheating allegation, which I don't think ANT is doing, it has to be doing some fortune telling/future reading to score that high at all.
If you look at the SWEBench official submissions: https://github.com/SWE-bench/experiments/tree/main/evaluatio..., filter all models after Sonnet 4, and aggregate ALL models' submission across 500 problems, what I found that the aggregated resolution rate is 93% (sharp).
Mythos gets 93.7%, meaning it solves problems that no other models could ever solve. I took a look at those problems, then I became even more suspicious, for the remaining 7% problems, it is almost impossible to resolve those issues without looking at the testing patch ahead of time, because how drastically the solution itself deviates from the problem statement, it almost feels like it is trying to solve a different problem.
Not that I am saying Mythos is cheating, but it might be too capable to remember all states of said repos, that it is able to reverse engineer the TRUE problem statement by diffing within its own internal memory. I think it could be a unique phenomena of evaluation awareness. Otherwise I genuinely couldn't think of exactly how it could be this precise in deciphering such unspecific problem statements.
The content encoder, the structure prior in this model is based on SongTi(SourceHans Serif), seal script would be foreign to it
Motivation behind the project is that I feel font generation has made a long way after zi2zi, but feels still not quite live up to my expectation where it can become a practical technology in creating fonts people can use.
zi2zi-JiT is thus created during last Christmas, aiming to create production grade CJK fonts that can be seen/used in everyday life, instead of simply being a research project.
So far, I have created 2 full set Chinese fonts using zi2zi-JiT, each with 6,763 Chinese characters created (GB2312 standard), from ancient Chinese books/caligraphies:
https://github.com/kaonashi-tyc/Zi-QuanHengDuLiang
https://github.com/kaonashi-tyc/Zi-XuanZongTi
All created in matters of 2-3 days, and free for commercial uses.
There are still some issue with current approach, but I am happy to hear you guys feedbacks and improve from here :)