TheBloke commited on
Commit
af2ad74
1 Parent(s): d17e02f

Update for Transformers GPTQ support

Browse files
README.md CHANGED
@@ -248,23 +248,26 @@ extra_gated_prompt: >-
248
  Please read the BigCode [OpenRAIL-M
249
  license](https://huggingface.co/spaces/bigcode/bigcode-model-license-agreement)
250
  agreement before accepting it.
251
-
252
  extra_gated_fields:
253
  I accept the above license agreement, and will use the Model complying with the set of use restrictions and sharing requirements: checkbox
254
  ---
255
 
256
  <!-- header start -->
257
- <div style="width: 100%;">
258
- <img src="https://i.imgur.com/EBdldam.jpg" alt="TheBlokeAI" style="width: 100%; min-width: 400px; display: block; margin: auto;">
 
259
  </div>
260
  <div style="display: flex; justify-content: space-between; width: 100%;">
261
  <div style="display: flex; flex-direction: column; align-items: flex-start;">
262
- <p><a href="https://discord.gg/Jq4vkcDakD">Chat & support: my new Discord server</a></p>
263
  </div>
264
  <div style="display: flex; flex-direction: column; align-items: flex-end;">
265
- <p><a href="https://www.patreon.com/TheBlokeAI">Want to contribute? TheBloke's Patreon page</a></p>
266
  </div>
267
  </div>
 
 
268
  <!-- header end -->
269
 
270
  # Bigcode's Starcoder GPTQ
@@ -281,9 +284,9 @@ It is the result of quantising to 4bit using [AutoGPTQ](https://github.com/PanQi
281
 
282
  ## Prompting
283
 
284
- The model was trained on GitHub code.
285
 
286
- As such it is _not_ an instruction model and commands like "Write a function that computes the square root." do not work well.
287
 
288
  However, by using the [Tech Assistant prompt](https://huggingface.co/datasets/bigcode/ta-prompt) you can turn it into a capable technical assistant.
289
 
@@ -349,11 +352,12 @@ It was created without group_size to lower VRAM requirements, and with --act-ord
349
  * Parameters: Groupsize = -1. Act Order / desc_act = True.
350
 
351
  <!-- footer start -->
 
352
  ## Discord
353
 
354
  For further support, and discussions on these models and AI in general, join us at:
355
 
356
- [TheBloke AI's Discord server](https://discord.gg/Jq4vkcDakD)
357
 
358
  ## Thanks, and how to contribute.
359
 
@@ -368,12 +372,15 @@ Donaters will get priority support on any and all AI/LLM/model questions and req
368
  * Patreon: https://patreon.com/TheBlokeAI
369
  * Ko-Fi: https://ko-fi.com/TheBlokeAI
370
 
371
- **Special thanks to**: Luke from CarbonQuill, Aemon Algiz, Dmitriy Samsonov.
 
 
372
 
373
- **Patreon special mentions**: Ajan Kanaga, Kalila, Derek Yates, Sean Connelly, Luke, Nathan LeClaire, Trenton Dambrowitz, Mano Prime, David Flickinger, vamX, Nikolai Manek, senxiiz, Khalefa Al-Ahmad, Illia Dulskyi, trip7s trip, Jonathan Leane, Talal Aujan, Artur Olbinski, Cory Kujawski, Joseph William Delisle, Pyrater, Oscar Rangel, Lone Striker, Luke Pendergrass, Eugene Pentland, Johann-Peter Hartmann.
374
 
375
  Thank you to all my generous patrons and donaters!
376
 
 
 
377
  <!-- footer end -->
378
 
379
  # Original model card: Bigcode's Starcoder
@@ -395,7 +402,7 @@ Play with the model on the [StarCoder Playground](https://huggingface.co/spaces/
395
 
396
  ## Model Summary
397
 
398
- The StarCoder models are 15.5B parameter models trained on 80+ programming languages from [The Stack (v1.2)](https://huggingface.co/datasets/bigcode/the-stack), with opt-out requests excluded. The model uses [Multi Query Attention](https://arxiv.org/abs/1911.02150), [a context window of 8192 tokens](https://arxiv.org/abs/2205.14135), and was trained using the [Fill-in-the-Middle objective](https://arxiv.org/abs/2207.14255) on 1 trillion tokens.
399
 
400
  - **Repository:** [bigcode/Megatron-LM](https://github.com/bigcode-project/Megatron-LM)
401
  - **Project Website:** [bigcode-project.org](https://www.bigcode-project.org)
@@ -444,7 +451,7 @@ The pretraining dataset of the model was filtered for permissive licenses only.
444
 
445
  # Limitations
446
 
447
- The model has been trained on source code from 80+ programming languages. The predominant natural language in source code is English although other languages are also present. As such the model is capable of generating code snippets provided some context but the generated code is not guaranteed to work as intended. It can be inefficient, contain bugs or exploits. See [the paper](https://drive.google.com/file/d/1cN-b9GnWtHzQRoE7M7gAEyivY0kl4BYs/view) for an in-depth discussion of the model limitations.
448
 
449
  # Training
450
 
@@ -471,7 +478,7 @@ The model is licensed under the BigCode OpenRAIL-M v1 license agreement. You can
471
  # Citation
472
  ```
473
  @article{li2023starcoder,
474
- title={StarCoder: may the source be with you!},
475
  author={Raymond Li and Loubna Ben Allal and Yangtian Zi and Niklas Muennighoff and Denis Kocetkov and Chenghao Mou and Marc Marone and Christopher Akiki and Jia Li and Jenny Chim and Qian Liu and Evgenii Zheltonozhskii and Terry Yue Zhuo and Thomas Wang and Olivier Dehaene and Mishig Davaadorj and Joel Lamy-Poirier and João Monteiro and Oleh Shliazhko and Nicolas Gontier and Nicholas Meade and Armel Zebaze and Ming-Ho Yee and Logesh Kumar Umapathi and Jian Zhu and Benjamin Lipkin and Muhtasham Oblokulov and Zhiruo Wang and Rudra Murthy and Jason Stillerman and Siva Sankalp Patel and Dmitry Abulkhanov and Marco Zocca and Manan Dey and Zhihan Zhang and Nour Fahmy and Urvashi Bhattacharyya and Wenhao Yu and Swayam Singh and Sasha Luccioni and Paulo Villegas and Maxim Kunakov and Fedor Zhdanov and Manuel Romero and Tony Lee and Nadav Timor and Jennifer Ding and Claire Schlesinger and Hailey Schoelkopf and Jan Ebert and Tri Dao and Mayank Mishra and Alex Gu and Jennifer Robinson and Carolyn Jane Anderson and Brendan Dolan-Gavitt and Danish Contractor and Siva Reddy and Daniel Fried and Dzmitry Bahdanau and Yacine Jernite and Carlos Muñoz Ferrandis and Sean Hughes and Thomas Wolf and Arjun Guha and Leandro von Werra and Harm de Vries},
476
  year={2023},
477
  eprint={2305.06161},
 
248
  Please read the BigCode [OpenRAIL-M
249
  license](https://huggingface.co/spaces/bigcode/bigcode-model-license-agreement)
250
  agreement before accepting it.
251
+
252
  extra_gated_fields:
253
  I accept the above license agreement, and will use the Model complying with the set of use restrictions and sharing requirements: checkbox
254
  ---
255
 
256
  <!-- header start -->
257
+ <!-- 200823 -->
258
+ <div style="width: auto; margin-left: auto; margin-right: auto">
259
+ <img src="https://i.imgur.com/EBdldam.jpg" alt="TheBlokeAI" style="width: 100%; min-width: 400px; display: block; margin: auto;">
260
  </div>
261
  <div style="display: flex; justify-content: space-between; width: 100%;">
262
  <div style="display: flex; flex-direction: column; align-items: flex-start;">
263
+ <p style="margin-top: 0.5em; margin-bottom: 0em;"><a href="https://discord.gg/theblokeai">Chat & support: TheBloke's Discord server</a></p>
264
  </div>
265
  <div style="display: flex; flex-direction: column; align-items: flex-end;">
266
+ <p style="margin-top: 0.5em; margin-bottom: 0em;"><a href="https://www.patreon.com/TheBlokeAI">Want to contribute? TheBloke's Patreon page</a></p>
267
  </div>
268
  </div>
269
+ <div style="text-align:center; margin-top: 0em; margin-bottom: 0em"><p style="margin-top: 0.25em; margin-bottom: 0em;">TheBloke's LLM work is generously supported by a grant from <a href="https://a16z.com">andreessen horowitz (a16z)</a></p></div>
270
+ <hr style="margin-top: 1.0em; margin-bottom: 1.0em;">
271
  <!-- header end -->
272
 
273
  # Bigcode's Starcoder GPTQ
 
284
 
285
  ## Prompting
286
 
287
+ The model was trained on GitHub code.
288
 
289
+ As such it is _not_ an instruction model and commands like "Write a function that computes the square root." do not work well.
290
 
291
  However, by using the [Tech Assistant prompt](https://huggingface.co/datasets/bigcode/ta-prompt) you can turn it into a capable technical assistant.
292
 
 
352
  * Parameters: Groupsize = -1. Act Order / desc_act = True.
353
 
354
  <!-- footer start -->
355
+ <!-- 200823 -->
356
  ## Discord
357
 
358
  For further support, and discussions on these models and AI in general, join us at:
359
 
360
+ [TheBloke AI's Discord server](https://discord.gg/theblokeai)
361
 
362
  ## Thanks, and how to contribute.
363
 
 
372
  * Patreon: https://patreon.com/TheBlokeAI
373
  * Ko-Fi: https://ko-fi.com/TheBlokeAI
374
 
375
+ **Special thanks to**: Aemon Algiz.
376
+
377
+ **Patreon special mentions**: Sam, theTransient, Jonathan Leane, Steven Wood, webtim, Johann-Peter Hartmann, Geoffrey Montalvo, Gabriel Tamborski, Willem Michiel, John Villwock, Derek Yates, Mesiah Bishop, Eugene Pentland, Pieter, Chadd, Stephen Murray, Daniel P. Andersen, terasurfer, Brandon Frisco, Thomas Belote, Sid, Nathan LeClaire, Magnesian, Alps Aficionado, Stanislav Ovsiannikov, Alex, Joseph William Delisle, Nikolai Manek, Michael Davis, Junyu Yang, K, J, Spencer Kim, Stefan Sabev, Olusegun Samson, transmissions 11, Michael Levine, Cory Kujawski, Rainer Wilmers, zynix, Kalila, Luke @flexchar, Ajan Kanaga, Mandus, vamX, Ai Maven, Mano Prime, Matthew Berman, subjectnull, Vitor Caleffi, Clay Pascal, biorpg, alfie_i, 阿明, Jeffrey Morgan, ya boyyy, Raymond Fosdick, knownsqashed, Olakabola, Leonard Tan, ReadyPlayerEmma, Enrico Ros, Dave, Talal Aujan, Illia Dulskyi, Sean Connelly, senxiiz, Artur Olbinski, Elle, Raven Klaugh, Fen Risland, Deep Realms, Imad Khwaja, Fred von Graf, Will Dee, usrbinkat, SuperWojo, Alexandros Triantafyllidis, Swaroop Kallakuri, Dan Guido, John Detwiler, Pedro Madruga, Iucharbius, Viktor Bowallius, Asp the Wyvern, Edmond Seymore, Trenton Dambrowitz, Space Cruiser, Spiking Neurons AB, Pyrater, LangChain4j, Tony Hughes, Kacper Wikieł, Rishabh Srivastava, David Ziegler, Luke Pendergrass, Andrey, Gabriel Puliatti, Lone Striker, Sebastain Graf, Pierre Kircher, Randy H, NimbleBox.ai, Vadim, danny, Deo Leter
378
 
 
379
 
380
  Thank you to all my generous patrons and donaters!
381
 
382
+ And thank you again to a16z for their generous grant.
383
+
384
  <!-- footer end -->
385
 
386
  # Original model card: Bigcode's Starcoder
 
402
 
403
  ## Model Summary
404
 
405
+ The StarCoder models are 15.5B parameter models trained on 80+ programming languages from [The Stack (v1.2)](https://huggingface.co/datasets/bigcode/the-stack), with opt-out requests excluded. The model uses [Multi Query Attention](https://arxiv.org/abs/1911.02150), [a context window of 8192 tokens](https://arxiv.org/abs/2205.14135), and was trained using the [Fill-in-the-Middle objective](https://arxiv.org/abs/2207.14255) on 1 trillion tokens.
406
 
407
  - **Repository:** [bigcode/Megatron-LM](https://github.com/bigcode-project/Megatron-LM)
408
  - **Project Website:** [bigcode-project.org](https://www.bigcode-project.org)
 
451
 
452
  # Limitations
453
 
454
+ The model has been trained on source code from 80+ programming languages. The predominant natural language in source code is English although other languages are also present. As such the model is capable of generating code snippets provided some context but the generated code is not guaranteed to work as intended. It can be inefficient, contain bugs or exploits. See [the paper](https://drive.google.com/file/d/1cN-b9GnWtHzQRoE7M7gAEyivY0kl4BYs/view) for an in-depth discussion of the model limitations.
455
 
456
  # Training
457
 
 
478
  # Citation
479
  ```
480
  @article{li2023starcoder,
481
+ title={StarCoder: may the source be with you!},
482
  author={Raymond Li and Loubna Ben Allal and Yangtian Zi and Niklas Muennighoff and Denis Kocetkov and Chenghao Mou and Marc Marone and Christopher Akiki and Jia Li and Jenny Chim and Qian Liu and Evgenii Zheltonozhskii and Terry Yue Zhuo and Thomas Wang and Olivier Dehaene and Mishig Davaadorj and Joel Lamy-Poirier and João Monteiro and Oleh Shliazhko and Nicolas Gontier and Nicholas Meade and Armel Zebaze and Ming-Ho Yee and Logesh Kumar Umapathi and Jian Zhu and Benjamin Lipkin and Muhtasham Oblokulov and Zhiruo Wang and Rudra Murthy and Jason Stillerman and Siva Sankalp Patel and Dmitry Abulkhanov and Marco Zocca and Manan Dey and Zhihan Zhang and Nour Fahmy and Urvashi Bhattacharyya and Wenhao Yu and Swayam Singh and Sasha Luccioni and Paulo Villegas and Maxim Kunakov and Fedor Zhdanov and Manuel Romero and Tony Lee and Nadav Timor and Jennifer Ding and Claire Schlesinger and Hailey Schoelkopf and Jan Ebert and Tri Dao and Mayank Mishra and Alex Gu and Jennifer Robinson and Carolyn Jane Anderson and Brendan Dolan-Gavitt and Danish Contractor and Siva Reddy and Daniel Fried and Dzmitry Bahdanau and Yacine Jernite and Carlos Muñoz Ferrandis and Sean Hughes and Thomas Wolf and Arjun Guha and Leandro von Werra and Harm de Vries},
483
  year={2023},
484
  eprint={2305.06161},
config.json CHANGED
@@ -1,39 +1,50 @@
1
  {
2
- "_name_or_path": "/fsx/bigcode/experiments/pretraining/conversions/starcoderpy/large-model",
3
- "activation_function": "gelu",
4
- "architectures": [
5
- "GPTBigCodeForCausalLM"
6
- ],
7
- "attention_softmax_in_fp32": true,
8
- "multi_query": true,
9
- "attn_pdrop": 0.1,
10
- "bos_token_id": 0,
11
- "embd_pdrop": 0.1,
12
- "eos_token_id": 0,
13
- "inference_runner": 0,
14
- "initializer_range": 0.02,
15
- "layer_norm_epsilon": 1e-05,
16
- "max_batch_size": null,
17
- "max_sequence_length": null,
18
- "model_type": "gpt_bigcode",
19
- "n_embd": 6144,
20
- "n_head": 48,
21
- "n_inner": 24576,
22
- "n_layer": 40,
23
- "n_positions": 8192,
24
- "pad_key_length": true,
25
- "pre_allocate_kv_cache": false,
26
- "resid_pdrop": 0.1,
27
- "scale_attention_softmax_in_fp32": true,
28
- "scale_attn_weights": true,
29
- "summary_activation": null,
30
- "summary_first_dropout": 0.1,
31
- "summary_proj_to_labels": true,
32
- "summary_type": "cls_index",
33
- "summary_use_proj": true,
34
- "torch_dtype": "float32",
35
- "transformers_version": "4.28.1",
36
- "use_cache": true,
37
- "validate_runner_input": true,
38
- "vocab_size": 49152
 
 
 
 
 
 
 
 
 
 
 
39
  }
 
1
  {
2
+ "_name_or_path": "/fsx/bigcode/experiments/pretraining/conversions/starcoderpy/large-model",
3
+ "activation_function": "gelu",
4
+ "architectures": [
5
+ "GPTBigCodeForCausalLM"
6
+ ],
7
+ "attention_softmax_in_fp32": true,
8
+ "multi_query": true,
9
+ "attn_pdrop": 0.1,
10
+ "bos_token_id": 0,
11
+ "embd_pdrop": 0.1,
12
+ "eos_token_id": 0,
13
+ "inference_runner": 0,
14
+ "initializer_range": 0.02,
15
+ "layer_norm_epsilon": 1e-05,
16
+ "max_batch_size": null,
17
+ "max_sequence_length": null,
18
+ "model_type": "gpt_bigcode",
19
+ "n_embd": 6144,
20
+ "n_head": 48,
21
+ "n_inner": 24576,
22
+ "n_layer": 40,
23
+ "n_positions": 8192,
24
+ "pad_key_length": true,
25
+ "pre_allocate_kv_cache": false,
26
+ "resid_pdrop": 0.1,
27
+ "scale_attention_softmax_in_fp32": true,
28
+ "scale_attn_weights": true,
29
+ "summary_activation": null,
30
+ "summary_first_dropout": 0.1,
31
+ "summary_proj_to_labels": true,
32
+ "summary_type": "cls_index",
33
+ "summary_use_proj": true,
34
+ "torch_dtype": "float32",
35
+ "transformers_version": "4.28.1",
36
+ "use_cache": true,
37
+ "validate_runner_input": true,
38
+ "vocab_size": 49152,
39
+ "quantization_config": {
40
+ "bits": 4,
41
+ "group_size": -1,
42
+ "damp_percent": 0.01,
43
+ "desc_act": true,
44
+ "sym": true,
45
+ "true_sequential": true,
46
+ "model_name_or_path": null,
47
+ "model_file_base_name": "model",
48
+ "quant_method": "gptq"
49
+ }
50
  }
gptq_model-4bit--1g.safetensors → model.safetensors RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:3ed6050a3c139c2abcd7e3187a2f2d49891748b2b229e79190bccb88932af39c
3
- size 8906589520
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d51843253aaf72e7116f64f879eab57b75bd5c03e6a9f15dcdc29bf9002e4201
3
+ size 8906589576
quantize_config.json CHANGED
@@ -6,5 +6,5 @@
6
  "sym": true,
7
  "true_sequential": true,
8
  "model_name_or_path": null,
9
- "model_file_base_name": null
10
  }
 
6
  "sym": true,
7
  "true_sequential": true,
8
  "model_name_or_path": null,
9
+ "model_file_base_name": "model"
10
  }