And yes totally. The other massive impact could be in source analysis. I have started using Mistral-Hermes for text annotation and it is both impressive and very fast.
What I'm saying is you review its findings and verify it. Use your own intelligence to determine if its bullshit or not.
A great use would be to enable one to have conversations with Pascal or Leibnitz, etc.
For instance, I published online the complete text of the Mémoires de Saint-Simon (written in 1745-1755, but describing the second part of the reign of Louis XIV and the Régence, 1695-1721).
Saint-Simon was described by his contemporaries to be one of the greats conversationalists of his time. It would be so cool to chat with him.
While I don't think Saint-Simon is included, a French colleague did a few try with it that turned out better than ChatGPT.
I'm currently working on an extended historical model for French (from 1000-2000) and maybe Saint-Simon memoirs will be included as well.
> the completely dataset here: https://huggingface.co/datasets/Pclanglais/MonadGPT
Classic French transcription seems to be lacking. In particular, "s" used to be printed in a manner very similar to "f", but they're really s.
For example this:
> ce qui augmentoit ſes craintesc'eſt que certe innocente Vierge ne parloit iamais d'autre choſe aux Domeſtiques que du lcge d'Orl'cans donnant à connoitre à la façon dont elle en difcouroit que fon inclination eſtoit toute aux armes
should be spelled like this:
> ce qui augmentoit ses craintes c'est que cette innocente Vierge ne parloit jamais d'autre chose aux Domestiques que du ?? d'Orléans donnant à connoître à la façon dont elle en discouroit que son inclination étoit (or estoit) toute aux armes
Maybe there should be some kind of dictionary step before fine-tuning?
approximately 180,000 inscriptions
What is possible is to use a larger learning rate but this will be a hard trade-off with conversational capacities. Fine tuning is currently based on original texts with a synthetic prompt. The issues that people have noticed (repetitions, not remembering what was in the prompt) will be more significant if the learning rate is higher.
Maybe a solution will be to provide two different variant of the same model, one less immersive and more workable, and the other more immersive and buggy.