dpf··on BigCode Project Releases StarCoder: A 15B Code LLMcode-davinci-002 is a base LM, and the other 3.5 models (text-davinci-{002,003}, gpt-3.5-turbo, and ChatGPT) use instruction tuning and/or RLHF. Source: https://platform.openai.com/docs/model-index-for-researchers
dpf··on On Chomsky and the Two Cultures of Statistical Learning (2011)Previous discussion:https://news.ycombinator.com/item?id=11951444 (2016)https://news.ycombinator.com/item?id=2591154 (2011)
dpf··on Roboschool: open-source software for robot simulationThese environments are often used as a testbed for reinforcement learning, e.g. https://arxiv.org/abs/1502.05477