Voltar ao ranking

Code for the paper "Language Models are Unsupervised Multitask Learners"

paper
Crescimento de estrelas
Estrelas
25k
Forks
5.9k
Crescimento semanal
Issues
148
10k20k
fev. de 2019jul. de 2021jan. de 2024jul. de 2026
ArtefatosPyPIpip install gpt-2
README

Status: Archive (code is provided as-is, no updates expected)

gpt-2

Code and models from the paper "Language Models are Unsupervised Multitask Learners".

You can read about GPT-2 and its staged release in our original blog post, 6 month follow-up post, and final post.

We have also released a dataset for researchers to study their behaviors.

* Note that our original parameter counts were wrong due to an error (in our previous blog posts and paper). Thus you may have seen small referred to as 117M and medium referred to as 345M.

Usage

This repository is meant to be a starting point for researchers and engineers to experiment with GPT-2.

For basic information, see our model card.

Some caveats

  • GPT-2 models' robustness and worst case behaviors are not well-understood. As with any machine-learned model, carefully evaluate GPT-2 for your use case, especially if used without fine-tuning or in safety-critical applications where reliability is important.
  • The dataset our GPT-2 models were trained on contains many texts with biases and factual inaccuracies, and thus GPT-2 models are likely to be biased and inaccurate as well.
  • To avoid having samples mistaken as human-written, we recommend clearly labeling samples as synthetic before wide dissemination. Our models are often incoherent or inaccurate in subtle ways, which takes more than a quick read for a human to notice.

Work with us

Please let us know if you’re doing interesting research with or working on applications of GPT-2! We’re especially interested in hearing from and potentially working with those who are studying

  • Potential malicious use cases and defenses against them (e.g. the detectability of synthetic text)
  • The extent of problematic content (e.g. bias) being baked into the models and effective mitigations

Development

See DEVELOPERS.md

Contributors

See CONTRIBUTORS.md

Citation

Please use the following bibtex entry:

@article{radford2019language,
  title={Language Models are Unsupervised Multitask Learners},
  author={Radford, Alec and Wu, Jeff and Child, Rewon and Luan, David and Amodei, Dario and Sutskever, Ilya},
  year={2019}
}

Future work

We may release code for evaluating the models on various benchmarks.

We are still considering release of the larger models.

License

Modified MIT

Repositórios relacionados
microsoft/qlib

Qlib is an AI-oriented Quant investment platform that aims to use AI tech to empower Quant Research, from exploring ideas to implementing productions. Qlib supports diverse ML modeling paradigms, including supervised learning, market dynamics modeling, and RL, and is now equipped with https://github.com/microsoft/RD-Agent to automate R&D process.

PythonPyPIMIT Licensequantitative-financemachine-learning
qlib.readthedocs.io/en/latest/
46.5k7.4k
mli/paper-reading

深度学习经典、新论文逐段精读

Apache License 2.0deep-learningpaper
33.6k2.8k
ipfs/ipfs

Peer-to-peer hypermedia protocol

MIT Licenseipfsp2p
ipfs.tech
23.1k1.5k
amusi/CVPR2026-Papers-with-Code

CVPR 2026 论文和开源项目合集

cvprcvpr2020
22.8k2.8k
kaixindelele/ChatPaper

Use ChatGPT to summarize the arXiv papers. 全流程加速科研,利用chatgpt进行论文全文总结+专业翻译+润色+审稿+审稿回复

PythonPyPIOtherarxivpaper
academic.chatpaper.top
19.7k1.9k
amusi/CVPR2025-Papers-with-Code

CVPR 2025 论文和开源项目合集

cvprcvpr2020
19.1k2.6k
amusi/CVPR2024-Papers-with-Code

CVPR 2024 论文和开源项目合集

cvprcvpr2020
18.7k2.6k
iamgio/quarkdown

🪐 Markdown with superpowers: from ideas to papers, presentations, websites, books, and knowledge bases.

KotlinGNU General Public License v3.0markdownmarkup-language
quarkdown.com
15.8k489
zziz/pwc

This repository is no longer maintained.

machine-learningpaper
15.3k2.4k
graykode/nlp-tutorial

Natural Language Processing Tutorial for Deep Learning Researchers

Jupyter NotebookMIT Licensenlpnatural-language-processing
reddit.com/r/MachineLearning/comments/amfinl/project_nlptutoral_repository_who_is_studying/
14.9k3.9k
jindongwang/transferlearning

Transfer learning / domain adaptation / domain generalization / multi-task learning etc. Papers, codes, datasets, applications, tutorials.-迁移学习

PythonPyPIMIT Licensetransferlearningdomain-adaptation
transferlearning.xyz
14.3k3.8k
PaperMC/Paper

The most widely used, high performance Minecraft server that aims to fix gameplay and mechanics inconsistencies

JavaMavenOtherminecraftbukkit
papermc.io/software/paper
12.5k3.5k