slimfrikha-tii commited on
Commit
c212cca
·
0 Parent(s):

falcon3 release

Browse files
.gitattributes ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,209 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ - fr
5
+ - es
6
+ - pt
7
+ tags:
8
+ - falcon3
9
+ license: other
10
+ license_name: falcon-llm-license
11
+ license_link: https://falconllm.tii.ae/falcon-terms-and-conditions.html
12
+ ---
13
+
14
+ <div align="center">
15
+ <img src="https://huggingface.co/datasets/tiiuae/documentation-images/resolve/main/general/falco3-logo.png" alt="drawing" width="500"/>
16
+ </div>
17
+
18
+ # Falcon3-1B-Base
19
+
20
+ **Falcon3** family of Open Foundation Models is a set of pretrained and instruct LLMs ranging from 1B to 10B parameters.
21
+
22
+ This repository contains the **Falcon3-1B-Base**. It achieves strong results on reasoning, language understanding, instruction following, code and mathematics tasks.
23
+ Falcon3-1B-Base supports 4 languages (English, French, Spanish, Portuguese) and a context length of up to 4K.
24
+ It was pruned in terms of depth, width, number of heads, and embedding channels from a larger 3B Falcon model, and was efficiently trained on only 80 GT using a knowledge distillation objective.
25
+
26
+ ⚠️ **This is a raw, pretrained model, which should be further finetuned using SFT, RLHF, continued pretraining, etc. for most use cases.**
27
+
28
+ ## Model Details
29
+ - Architecture
30
+ - Transformer-based causal decoder-only architecture
31
+ - 18 decoder blocks
32
+ - Grouped Query Attention (GQA) for faster inference: 8 query heads and 4 key-value heads
33
+ - Wider head dimension: 256
34
+ - High RoPE value to support long context understanding: 1000042
35
+ - Uses SwiGLU and RMSNorm
36
+ - 4K context length
37
+ - 131K vocab size
38
+ - Pruned and healed using larger Falcon models (3B and 7B respectively) on only 80 Gigatokens of datasets comprising of web, code, STEM, high quality and multilingual data using 256 H100 GPU chips
39
+ - Supports EN, FR, ES, PT
40
+ - Developed by [Technology Innovation Institute](https://www.tii.ae)
41
+ - License: TII Falcon-LLM License 2.0
42
+ - Model Release Date: December 2024
43
+
44
+
45
+ ## Getting started
46
+
47
+ <details>
48
+ <summary> Click to expand </summary>
49
+
50
+ ```python
51
+ import torch
52
+ from transformers import pipeline
53
+
54
+ pipe = pipeline(
55
+ "text-generation",
56
+ model="tiiuae/Falcon3-1B-Base",
57
+ torch_dtype=torch.bfloat16,
58
+ device_map="auto"
59
+ )
60
+ response = pipe("Question: How many hours in one day? Answer: ")
61
+ print(response[0]['generated_text'])
62
+ ```
63
+
64
+ </details>
65
+
66
+ <br>
67
+
68
+ ## Benchmarks
69
+ We report in the following table our internal pipeline benchmarks:
70
+
71
+
72
+
73
+ <table border="1" style="width: 100%; text-align: center; border-collapse: collapse;">
74
+ <colgroup>
75
+ <col style="width: 10%;">
76
+ <col style="width: 10%;">
77
+ <col style="width: 7%;">
78
+ <col style="width: 7%;">
79
+ <col style="width: 7%;">
80
+ <col style="background-color: rgba(80, 15, 213, 0.5); width: 7%;">
81
+ </colgroup>
82
+ <thead>
83
+ <tr>
84
+ <th>Category</th>
85
+ <th>Benchmark</th>
86
+ <th>Llama-3.2-1B</th>
87
+ <th>Qwen2.5-1.5B</th>
88
+ <th>SmolLM2-1.7B</th>
89
+ <th>Falcon3-1B-Base</th>
90
+ </tr>
91
+ </thead>
92
+ <tbody>
93
+ <tr>
94
+ <td rowspan="3">General</td>
95
+ <td>MMLU (5-shot)</td>
96
+ <td>31.1</td>
97
+ <td><b>61.0</b></td>
98
+ <td>50.1</td>
99
+ <td>42.5</td>
100
+ </tr>
101
+ <tr>
102
+ <td>MMLU-PRO (5-shot)</td>
103
+ <td>11.7</td>
104
+ <td><b>28.4</b></td>
105
+ <td>21.3</td>
106
+ <td>16.1</td>
107
+ </tr>
108
+ <tr>
109
+ <td>IFEval</td>
110
+ <td>14.8</td>
111
+ <td><b>26.0</b></td>
112
+ <td>24.2</td>
113
+ <td>25.2</td>
114
+ </tr>
115
+ <tr>
116
+ <td rowspan="2">Math</td>
117
+ <td>GSM8K (5-shot)</td>
118
+ <td>6.6</td>
119
+ <td><b>62.2</b></td>
120
+ <td>31.0</td>
121
+ <td>34.3</td>
122
+ </tr>
123
+ <tr>
124
+ <td>MATH Lvl-5 (4-shot)</td>
125
+ <td>0.2</td>
126
+ <td><b>6.7</b></td>
127
+ <td>1.4</td>
128
+ <td>2.2</td>
129
+ </tr>
130
+ <tr>
131
+ <td rowspan="4">Reasoning</td>
132
+ <td>Arc Challenge (25-shot)</td>
133
+ <td>40.2</td>
134
+ <td><b>54.8</b></td>
135
+ <td>54.1</td>
136
+ <td>48.1</td>
137
+ </tr>
138
+ <tr>
139
+ <td>GPQA (0-shot)</td>
140
+ <td>24.2</td>
141
+ <td>28.1</td>
142
+ <td><b>28.9</b></td>
143
+ <td>28.1</td>
144
+ </tr>
145
+ <tr>
146
+ <td>MUSR (0-shot)</td>
147
+ <td>34.5</td>
148
+ <td>35.5</td>
149
+ <td>34.7</td>
150
+ <td><b>41.9</b></td>
151
+ </tr>
152
+ <tr>
153
+ <td>BBH (3-shot)</td>
154
+ <td>31.2</td>
155
+ <td><b>41.1</b></td>
156
+ <td>34.2</td>
157
+ <td>36.0</td>
158
+ </tr>
159
+ <tr>
160
+ <td rowspan="4">CommonSense Understanding</td>
161
+ <td>PIQA (0-shot)</td>
162
+ <td>74.5</td>
163
+ <td>76.0</td>
164
+ <td><b>77.5</b></td>
165
+ <td>74.5</td>
166
+ </tr>
167
+ <tr>
168
+ <td>SciQ (0-shot)</td>
169
+ <td>88.5</td>
170
+ <td><b>93.1</b></td>
171
+ <td>90.8</td>
172
+ <td>91.1</td>
173
+ </tr>
174
+ <tr>
175
+ <td>Winogrande (0-shot)</td>
176
+ <td>60.4</td>
177
+ <td>63.0</td>
178
+ <td><b>66.1</b></td>
179
+ <td>61.2</td>
180
+ </tr>
181
+ <tr>
182
+ <td>OpenbookQA (0-shot)</td>
183
+ <td>37.4</td>
184
+ <td>40.4</td>
185
+ <td><b>44.0</b></td>
186
+ <td>41.0</td>
187
+ </tr>
188
+ </tbody>
189
+ </table>
190
+
191
+ ## Useful links
192
+ - View our [release blogpost](https://huggingface.co/blog/falcon3).
193
+ - Feel free to join [our discord server](https://discord.gg/fwXpMyGc) if you have any questions or to interact with our researchers and developers.
194
+
195
+ ## Technical Report
196
+ Coming soon....
197
+
198
+ ## Citation
199
+ If the Falcon3 family of models were helpful to your work, feel free to give us a cite.
200
+
201
+ ```
202
+ @misc{Falcon3,
203
+ title = {The Falcon 3 Family of Open Models},
204
+ url = {https://huggingface.co/blog/falcon3},
205
+ author = {Falcon-LLM Team},
206
+ month = {December},
207
+ year = {2024}
208
+ }
209
+ ```
config.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "LlamaForCausalLM"
4
+ ],
5
+ "activation": "swiglu",
6
+ "attention_bias": false,
7
+ "attention_dropout": 0.0,
8
+ "eos_token_id": 11,
9
+ "head_dim": 256,
10
+ "hidden_act": "silu",
11
+ "hidden_size": 2048,
12
+ "intermediate_size": 8192,
13
+ "max_position_embeddings": 4096,
14
+ "mlp_bias": false,
15
+ "model_type": "llama",
16
+ "num_attention_heads": 8,
17
+ "num_hidden_layers": 18,
18
+ "num_key_value_heads": 4,
19
+ "pretraining_tp": 1,
20
+ "rms_norm_eps": 1e-06,
21
+ "rope_scaling": null,
22
+ "rope_theta": 1000042,
23
+ "tie_word_embeddings": false,
24
+ "torch_dtype": "bfloat16",
25
+ "transformers_version": "4.46.1",
26
+ "use_cache": true,
27
+ "vocab_size": 131072
28
+ }
generation_config.json ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "eos_token_id": 11,
4
+ "transformers_version": "4.46.1"
5
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9c2c3117923ab2e22e8d07a92cd8fd069ce70a8b1c8557894194bbfd8fcee196
3
+ size 3338836632
model.safetensors.index.json ADDED
@@ -0,0 +1,172 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metadata": {
3
+ "total_size": 6677635072
4
+ },
5
+ "weight_map": {
6
+ "lm_head.weight": "model-00002-of-00002.safetensors",
7
+ "model.embed_tokens.weight": "model-00001-of-00002.safetensors",
8
+ "model.layers.0.input_layernorm.weight": "model-00001-of-00002.safetensors",
9
+ "model.layers.0.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
10
+ "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
11
+ "model.layers.0.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
12
+ "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
13
+ "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
14
+ "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
15
+ "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
16
+ "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
17
+ "model.layers.1.input_layernorm.weight": "model-00001-of-00002.safetensors",
18
+ "model.layers.1.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
19
+ "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
20
+ "model.layers.1.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
21
+ "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
22
+ "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
23
+ "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
24
+ "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
25
+ "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
26
+ "model.layers.10.input_layernorm.weight": "model-00001-of-00002.safetensors",
27
+ "model.layers.10.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
28
+ "model.layers.10.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
29
+ "model.layers.10.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
30
+ "model.layers.10.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
31
+ "model.layers.10.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
32
+ "model.layers.10.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
33
+ "model.layers.10.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
34
+ "model.layers.10.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
35
+ "model.layers.11.input_layernorm.weight": "model-00001-of-00002.safetensors",
36
+ "model.layers.11.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
37
+ "model.layers.11.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
38
+ "model.layers.11.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
39
+ "model.layers.11.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
40
+ "model.layers.11.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
41
+ "model.layers.11.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
42
+ "model.layers.11.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
43
+ "model.layers.11.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
44
+ "model.layers.12.input_layernorm.weight": "model-00001-of-00002.safetensors",
45
+ "model.layers.12.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
46
+ "model.layers.12.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
47
+ "model.layers.12.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
48
+ "model.layers.12.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
49
+ "model.layers.12.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
50
+ "model.layers.12.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
51
+ "model.layers.12.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
52
+ "model.layers.12.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
53
+ "model.layers.13.input_layernorm.weight": "model-00001-of-00002.safetensors",
54
+ "model.layers.13.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
55
+ "model.layers.13.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
56
+ "model.layers.13.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
57
+ "model.layers.13.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
58
+ "model.layers.13.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
59
+ "model.layers.13.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
60
+ "model.layers.13.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
61
+ "model.layers.13.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
62
+ "model.layers.14.input_layernorm.weight": "model-00001-of-00002.safetensors",
63
+ "model.layers.14.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
64
+ "model.layers.14.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
65
+ "model.layers.14.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
66
+ "model.layers.14.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
67
+ "model.layers.14.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
68
+ "model.layers.14.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
69
+ "model.layers.14.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
70
+ "model.layers.14.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
71
+ "model.layers.15.input_layernorm.weight": "model-00002-of-00002.safetensors",
72
+ "model.layers.15.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
73
+ "model.layers.15.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
74
+ "model.layers.15.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
75
+ "model.layers.15.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
76
+ "model.layers.15.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
77
+ "model.layers.15.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
78
+ "model.layers.15.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
79
+ "model.layers.15.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
80
+ "model.layers.16.input_layernorm.weight": "model-00002-of-00002.safetensors",
81
+ "model.layers.16.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
82
+ "model.layers.16.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
83
+ "model.layers.16.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
84
+ "model.layers.16.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
85
+ "model.layers.16.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
86
+ "model.layers.16.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
87
+ "model.layers.16.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
88
+ "model.layers.16.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
89
+ "model.layers.17.input_layernorm.weight": "model-00002-of-00002.safetensors",
90
+ "model.layers.17.mlp.down_proj.weight": "model-00002-of-00002.safetensors",
91
+ "model.layers.17.mlp.gate_proj.weight": "model-00002-of-00002.safetensors",
92
+ "model.layers.17.mlp.up_proj.weight": "model-00002-of-00002.safetensors",
93
+ "model.layers.17.post_attention_layernorm.weight": "model-00002-of-00002.safetensors",
94
+ "model.layers.17.self_attn.k_proj.weight": "model-00002-of-00002.safetensors",
95
+ "model.layers.17.self_attn.o_proj.weight": "model-00002-of-00002.safetensors",
96
+ "model.layers.17.self_attn.q_proj.weight": "model-00002-of-00002.safetensors",
97
+ "model.layers.17.self_attn.v_proj.weight": "model-00002-of-00002.safetensors",
98
+ "model.layers.2.input_layernorm.weight": "model-00001-of-00002.safetensors",
99
+ "model.layers.2.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
100
+ "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
101
+ "model.layers.2.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
102
+ "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
103
+ "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
104
+ "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
105
+ "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
106
+ "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
107
+ "model.layers.3.input_layernorm.weight": "model-00001-of-00002.safetensors",
108
+ "model.layers.3.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
109
+ "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
110
+ "model.layers.3.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
111
+ "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
112
+ "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
113
+ "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
114
+ "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
115
+ "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
116
+ "model.layers.4.input_layernorm.weight": "model-00001-of-00002.safetensors",
117
+ "model.layers.4.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
118
+ "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
119
+ "model.layers.4.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
120
+ "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
121
+ "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
122
+ "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
123
+ "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
124
+ "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
125
+ "model.layers.5.input_layernorm.weight": "model-00001-of-00002.safetensors",
126
+ "model.layers.5.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
127
+ "model.layers.5.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
128
+ "model.layers.5.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
129
+ "model.layers.5.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
130
+ "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
131
+ "model.layers.5.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
132
+ "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
133
+ "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
134
+ "model.layers.6.input_layernorm.weight": "model-00001-of-00002.safetensors",
135
+ "model.layers.6.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
136
+ "model.layers.6.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
137
+ "model.layers.6.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
138
+ "model.layers.6.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
139
+ "model.layers.6.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
140
+ "model.layers.6.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
141
+ "model.layers.6.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
142
+ "model.layers.6.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
143
+ "model.layers.7.input_layernorm.weight": "model-00001-of-00002.safetensors",
144
+ "model.layers.7.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
145
+ "model.layers.7.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
146
+ "model.layers.7.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
147
+ "model.layers.7.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
148
+ "model.layers.7.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
149
+ "model.layers.7.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
150
+ "model.layers.7.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
151
+ "model.layers.7.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
152
+ "model.layers.8.input_layernorm.weight": "model-00001-of-00002.safetensors",
153
+ "model.layers.8.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
154
+ "model.layers.8.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
155
+ "model.layers.8.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
156
+ "model.layers.8.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
157
+ "model.layers.8.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
158
+ "model.layers.8.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
159
+ "model.layers.8.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
160
+ "model.layers.8.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
161
+ "model.layers.9.input_layernorm.weight": "model-00001-of-00002.safetensors",
162
+ "model.layers.9.mlp.down_proj.weight": "model-00001-of-00002.safetensors",
163
+ "model.layers.9.mlp.gate_proj.weight": "model-00001-of-00002.safetensors",
164
+ "model.layers.9.mlp.up_proj.weight": "model-00001-of-00002.safetensors",
165
+ "model.layers.9.post_attention_layernorm.weight": "model-00001-of-00002.safetensors",
166
+ "model.layers.9.self_attn.k_proj.weight": "model-00001-of-00002.safetensors",
167
+ "model.layers.9.self_attn.o_proj.weight": "model-00001-of-00002.safetensors",
168
+ "model.layers.9.self_attn.q_proj.weight": "model-00001-of-00002.safetensors",
169
+ "model.layers.9.self_attn.v_proj.weight": "model-00001-of-00002.safetensors",
170
+ "model.norm.weight": "model-00002-of-00002.safetensors"
171
+ }
172
+ }
special_tokens_map.json ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "additional_special_tokens": [
3
+ ">>TITLE<<",
4
+ ">>ABSTRACT<<",
5
+ ">>INTRODUCTION<<",
6
+ ">>SUMMARY<<",
7
+ ">>COMMENT<<",
8
+ ">>ANSWER<<",
9
+ ">>QUESTION<<",
10
+ ">>DOMAIN<<",
11
+ ">>EMAIL_ADDRESS<<",
12
+ ">>IP_ADDRESS<<",
13
+ "<|startoftext|>",
14
+ ">>IP_ADDRESS_0<<",
15
+ ">>IP_ADDRESS_1<<",
16
+ ">>IP_ADDRESS_2<<",
17
+ ">>IP_ADDRESS_3<<",
18
+ ">>IP_ADDRESS_4<<",
19
+ ">>IP_ADDRESS_5<<",
20
+ ">>IP_ADDRESS_6<<",
21
+ ">>IP_ADDRESS_7<<",
22
+ ">>IP_ADDRESS_8<<",
23
+ ">>IP_ADDRESS_9<<",
24
+ ">>PASSWORD<<",
25
+ ">>KEY<<"
26
+ ],
27
+ "eos_token": {
28
+ "content": "<|endoftext|>",
29
+ "lstrip": false,
30
+ "normalized": false,
31
+ "rstrip": false,
32
+ "single_word": false
33
+ },
34
+ "pad_token": {
35
+ "content": "<|pad|>",
36
+ "lstrip": false,
37
+ "normalized": false,
38
+ "rstrip": false,
39
+ "single_word": false
40
+ }
41
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
The diff for this file is too large to render. See raw diff