Text2Text Generation
Transformers
Safetensors
English
mistral
text-generation
text-generation-inference
Inference Endpoints
seungone commited on
Commit
218c3d1
1 Parent(s): 716839b

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +166 -41
README.md CHANGED
@@ -1,56 +1,181 @@
1
  ---
2
  tags:
3
- - merge
4
- - mergekit
5
- - lazymergekit
6
- - kaist-ai/prometheus-7b-v1.5-beta-1
7
- - kaist-ai/prometheus-7b-v1.5-beta-2
8
- base_model:
9
- - kaist-ai/prometheus-7b-v1.5-beta-1
10
- - kaist-ai/prometheus-7b-v1.5-beta-2
 
 
 
 
 
 
11
  ---
 
12
 
13
- # kaist-ai/prometheus-7b-v1.5-beta-merged
 
 
 
14
 
15
- kaist-ai/prometheus-7b-v1.5-beta-merged is a merge of the following models using [LazyMergekit](https://colab.research.google.com/drive/1obulZ1ROXHjYLn6PPZJwRR6GzgQogxxb?usp=sharing):
16
- * [kaist-ai/prometheus-7b-v1.5-beta-1](https://huggingface.co/kaist-ai/prometheus-7b-v1.5-beta-1)
17
- * [kaist-ai/prometheus-7b-v1.5-beta-2](https://huggingface.co/kaist-ai/prometheus-7b-v1.5-beta-2)
18
 
19
- ## 🧩 Configuration
 
 
 
20
 
21
- ```yaml
22
- models:
23
- - model: kaist-ai/prometheus-7b-v1.5-beta-1
24
- parameters:
25
- weight: 1.0
26
- - model: kaist-ai/prometheus-7b-v1.5-beta-2
27
- parameters:
28
- weight: 1.0
29
- merge_method: linear
30
- dtype: bfloat16
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
31
  ```
 
 
 
 
 
 
 
 
 
32
 
33
- ## 💻 Usage
 
34
 
35
- ```python
36
- !pip install -qU transformers accelerate
37
 
38
- from transformers import AutoTokenizer
39
- import transformers
40
- import torch
 
 
 
 
41
 
42
- model = "kaist-ai/kaist-ai/prometheus-7b-v1.5-beta-merged"
43
- messages = [{"role": "user", "content": "What is a large language model?"}]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
- tokenizer = AutoTokenizer.from_pretrained(model)
46
- prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
47
- pipeline = transformers.pipeline(
48
- "text-generation",
49
- model=model,
50
- torch_dtype=torch.float16,
51
- device_map="auto",
52
- )
53
 
54
- outputs = pipeline(prompt, max_new_tokens=256, do_sample=True, temperature=0.7, top_k=50, top_p=0.95)
55
- print(outputs[0]["generated_text"])
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
56
  ```
 
1
  ---
2
  tags:
3
+ - text2text-generation
4
+ datasets:
5
+ - prometheus-eval/Feedback-Collection
6
+ - prometheus-eval/Preference-Collection
7
+ license: apache-2.0
8
+ language:
9
+ - en
10
+ pipeline_tag: text2text-generation
11
+ library_name: transformers
12
+ metrics:
13
+ - pearsonr
14
+ - spearmanr
15
+ - kendall-tau
16
+ - accuracy
17
  ---
18
+ ## Links for Reference
19
 
20
+ - **Homepage: In Progress**
21
+ - **Repository:https://github.com/prometheus-eval/prometheus-eval**
22
+ - **Paper:https://arxiv.org/abs/2405.01535**
23
+ - **Point of Contact:[email protected]**
24
 
25
+ # TL;DR
26
+ Prometheus 2 is an alternative of GPT-4 evaluation when doing fine-grained evaluation of an underlying LLM & a Reward model for Reinforcement Learning from Human Feedback (RLHF).
27
+ ![plot](./finegrained_eval.JPG)
28
 
29
+ Prometheus 2 is a language model using [Mistral-Instruct](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2) as a base model.
30
+ It is fine-tuned on 100K feedback within the [Feedback Collection](https://huggingface.co/datasets/prometheus-eval/Feedback-Collection) and 200K feedback within the [Preference Collection](https://huggingface.co/datasets/prometheus-eval/Preference-Collection).
31
+ It is also made by weight merging to support both absolute grading (direct assessment) and relative grading (pairwise ranking).
32
+ The surprising thing is that we find weight merging also improves performance on each format.
33
 
34
+ # Model Details
35
+
36
+ ## Model Description
37
+
38
+ - **Model type:** Language model
39
+ - **Language(s) (NLP):** English
40
+ - **License:** Apache 2.0
41
+ - **Related Models:** [All Prometheus Checkpoints](https://huggingface.co/models?search=prometheus-eval/Prometheus)
42
+ - **Resources for more information:**
43
+ - [Research paper](https://arxiv.org/abs/2405.01535)
44
+ - [GitHub Repo](https://github.com/prometheus-eval/prometheus-eval)
45
+
46
+
47
+ Prometheus is trained with two different sizes (7B and 8x7B).
48
+ You could check the 8x7B sized LM on [this page](https://huggingface.co/prometheus-eval/prometheus-2-8x7b-v2.0).
49
+ Also, check out our dataset as well on [this page](https://huggingface.co/datasets/prometheus-eval/Feedback-Collection) and [this page](https://huggingface.co/datasets/prometheus-eval/Preference-Collection).
50
+
51
+ ## Prompt Format
52
+
53
+ We have made wrapper functions and classes to conveniently use Prometheus 2 at [our github repository](https://github.com/prometheus-eval/prometheus-eval).
54
+ We highly recommend you use it!
55
+
56
+ However, if you just want to use the model for your use case, please refer to the prompt format below.
57
+ Note that absolute grading and relative grading requires different prompt templates and system prompts.
58
+
59
+ ### Absolute Grading (Direct Assessment)
60
+ Prometheus requires 4 components in the input: An instruction, a response to evaluate, a score rubric, and a reference answer. You could refer to the prompt format below.
61
+ You should fill in the instruction, response, reference answer, criteria description, and score description for score in range of 1 to 5.
62
+
63
+ Fix the components with \{text\} inside.
64
  ```
65
+ ###Task Description:
66
+ An instruction (might include an Input inside it), a response to evaluate, a reference answer that gets a score of 5, and a score rubric representing a evaluation criteria are given.
67
+ 1. Write a detailed feedback that assess the quality of the response strictly based on the given score rubric, not evaluating in general.
68
+ 2. After writing a feedback, write a score that is an integer between 1 and 5. You should refer to the score rubric.
69
+ 3. The output format should look as follows: \"Feedback: (write a feedback for criteria) [RESULT] (an integer number between 1 and 5)\"
70
+ 4. Please do not generate any other opening, closing, and explanations.
71
+
72
+ ###The instruction to evaluate:
73
+ {orig_instruction}
74
 
75
+ ###Response to evaluate:
76
+ {orig_response}
77
 
78
+ ###Reference Answer (Score 5):
79
+ {orig_reference_answer}
80
 
81
+ ###Score Rubrics:
82
+ [{orig_criteria}]
83
+ Score 1: {orig_score1_description}
84
+ Score 2: {orig_score2_description}
85
+ Score 3: {orig_score3_description}
86
+ Score 4: {orig_score4_description}
87
+ Score 5: {orig_score5_description}
88
 
89
+ ###Feedback:
90
+ ```
91
+
92
+ After this, you should apply the conversation template of Mistral (not applying it might lead to unexpected behaviors).
93
+ You can find the conversation class at this [link](https://github.com/lm-sys/FastChat/blob/main/fastchat/conversation.py).
94
+ ```
95
+ conv = get_conv_template("mistral")
96
+ conv.set_system_message("You are a fair judge assistant tasked with providing clear, objective feedback based on specific criteria, ensuring each assessment reflects the absolute standards set for performance.")
97
+ conv.append_message(conv.roles[0], dialogs['instruction'])
98
+ conv.append_message(conv.roles[1], None)
99
+ prompt = conv.get_prompt()
100
+
101
+ x = tokenizer(prompt,truncation=False)
102
+ ```
103
+
104
+ As a result, a feedback and score decision will be generated, divided by a separating phrase ```[RESULT]```
105
+
106
+ ### Relative Grading (Pairwise Ranking)
107
+ Prometheus requires 4 components in the input: An instruction, 2 responses to evaluate, a score rubric, and a reference answer. You could refer to the prompt format below.
108
+ You should fill in the instruction, 2 responses, reference answer, and criteria description.
109
+
110
+ Fix the components with \{text\} inside.
111
+ ```
112
+ ###Task Description:
113
+ An instruction (might include an Input inside it), a response to evaluate, and a score rubric representing a evaluation criteria are given.
114
+ 1. Write a detailed feedback that assess the quality of two responses strictly based on the given score rubric, not evaluating in general.
115
+ 2. After writing a feedback, choose a better response between Response A and Response B. You should refer to the score rubric.
116
+ 3. The output format should look as follows: "Feedback: (write a feedback for criteria) [RESULT] (A or B)"
117
+ 4. Please do not generate any other opening, closing, and explanations.
118
 
119
+ ###Instruction:
120
+ {orig_instruction}
 
 
 
 
 
 
121
 
122
+ ###Response A:
123
+ {orig_response_A}
124
+
125
+ ###Response B:
126
+ {orig_response_B}
127
+
128
+ ###Reference Answer:
129
+ {orig_reference_answer}
130
+
131
+ ###Score Rubric:
132
+ {orig_criteria}
133
+
134
+ ###Feedback:
135
+ ```
136
+
137
+ After this, you should apply the conversation template of Mistral (not applying it might lead to unexpected behaviors).
138
+ You can find the conversation class at this [link](https://github.com/lm-sys/FastChat/blob/main/fastchat/conversation.py).
139
+ ```
140
+ conv = get_conv_template("mistral")
141
+ conv.set_system_message("You are a fair judge assistant assigned to deliver insightful feedback that compares individual performances, highlighting how each stands relative to others within the same cohort.")
142
+ conv.append_message(conv.roles[0], dialogs['instruction'])
143
+ conv.append_message(conv.roles[1], None)
144
+ prompt = conv.get_prompt()
145
+
146
+ x = tokenizer(prompt,truncation=False)
147
+ ```
148
+
149
+ As a result, a feedback and score decision will be generated, divided by a separating phrase ```[RESULT]```
150
+
151
+ ## License
152
+ Feedback Collection, Preference Collection, and Prometheus 2 are subject to OpenAI's Terms of Use for the generated data. If you suspect any violations, please reach out to us.
153
+
154
+
155
+ # Citation
156
+
157
+
158
+ If you find the following model helpful, please consider citing our paper!
159
+
160
+ **BibTeX:**
161
+
162
+ ```bibtex
163
+ @misc{kim2023prometheus,
164
+ title={Prometheus: Inducing Fine-grained Evaluation Capability in Language Models},
165
+ author={Seungone Kim and Jamin Shin and Yejin Cho and Joel Jang and Shayne Longpre and Hwaran Lee and Sangdoo Yun and Seongjin Shin and Sungdong Kim and James Thorne and Minjoon Seo},
166
+ year={2023},
167
+ eprint={2310.08491},
168
+ archivePrefix={arXiv},
169
+ primaryClass={cs.CL}
170
+ }
171
+ ```
172
+ ```bibtex
173
+ @misc{kim2024prometheus,
174
+ title={Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models},
175
+ author={Seungone Kim and Juyoung Suk and Shayne Longpre and Bill Yuchen Lin and Jamin Shin and Sean Welleck and Graham Neubig and Moontae Lee and Kyungjae Lee and Minjoon Seo},
176
+ year={2024},
177
+ eprint={2405.01535},
178
+ archivePrefix={arXiv},
179
+ primaryClass={cs.CL}
180
+ }
181
  ```