argilla
/

notus-7b-v1

@@ -1,43 +1,95 @@
 ---
-license: apache-2.0
-base_model: alignment-handbook/zephyr-7b-sft-full
-tags:
-- generated_from_trainer
 model-index:
-- name: notus-7b-dpo
   results: []
 ---
-<!-- This model card has been generated automatically according to the information the Trainer had access to. You
-should probably proofread and complete it, then remove this comment. -->
 # notus-7b-dpo
-This model is a fine-tuned version of [alignment-handbook/zephyr-7b-sft-full](https://huggingface.co/alignment-handbook/zephyr-7b-sft-full) on the None dataset.
-It achieves the following results on the evaluation set:
-- Loss: 0.4730
-- Rewards/chosen: -3.5289
-- Rewards/rejected: -7.3700
-- Rewards/accuracies: 0.8016
-- Rewards/margins: 3.8412
-- Logps/rejected: -316.3751
-- Logps/chosen: -334.3053
-- Logits/rejected: -2.1644
-- Logits/chosen: -2.4556
-## Model description
-More information needed
-## Intended uses & limitations
-More information needed
-## Training and evaluation data
-More information needed
-## Training procedure
 ### Training hyperparameters
@@ -88,10 +140,96 @@ The following hyperparameters were used during training:
 | 0.0059        | 2.8   | 2700 | 0.4694          | -3.4307        | -7.2484          | 0.7976             | 3.8177          | -315.1584      | -333.3234    | -2.1572         | -2.4483       |
 | 0.0054        | 2.91  | 2800 | 0.4707          | -3.4959        | -7.3283          | 0.8056             | 3.8324          | -315.9576      | -333.9758    | -2.1575         | -2.4491       |
 ### Framework versions
 - Transformers 4.35.0
 - Pytorch 2.1.1+cu121
 - Datasets 2.14.6
 - Tokenizers 0.14.1

 ---
 model-index:
+- name: notus-7b-dpo-lora
   results: []
+datasets:
+- argilla/ultrafeedback-binarized-avg-rating-for-dpo
+language:
+- en
+base_model: alignment-handbook/zephyr-7b-sft-full
+library_name: transformers
+pipeline_tag: text-generation
+tags:
+- dpo
+- preference
+- ultrafeedback
+license: apache-2.0
 ---
+# Model Card for Notus 7B
+Notus is going to be a collection of fine-tuned models using DPO, similarly to Zephyr, but mainly focused
+on the Direct Preference Optimization (DPO) step, aiming to incorporate preference feedback into the LLMs
+when fine-tuning those. Notus models are intended to be used as assistants via chat-like applications, and
+are evaluated with the MT-Bench and AlpacaEval benchmarks, to be directly compared with Zephyr fine-tuned models
+also using DPO.
+## Model Details
 # notus-7b-dpo
+### Model Description
+- **Developed by:** Argilla, Inc. (based on HuggingFace H4 and MistralAI previous efforts and amazing work)
+- **Shared by:** Argilla, Inc.
+- **Model type:** GPT-like 7B model DPO fine-tuned
+- **Language(s) (NLP):** Mainly English
+- **License:** Apache 2.0 (same as Zephyr 7B SFT and Mistral 7B v0.1)
+- **Finetuned from model:** [`alignment-handbook/zephyr-7b-sft-full`](https://huggingface.co/alignment-handbook/zephyr-7b-sft-full)
+### Model Sources [optional]
+- **Repository:** https://github.com/argilla-io/notus-7b-dpo
+- **Paper:** N/A
+- **Demo:** https://argilla-notus-chat-ui.hf.space/
+## Uses
+<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
+### Direct Use
+<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
+[More Information Needed]
+### Downstream Use [optional]
+<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
+[More Information Needed]
+### Out-of-Scope Use
+<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
+[More Information Needed]
+## Bias, Risks, and Limitations
+<!-- This section is meant to convey both technical and sociotechnical limitations. -->
+[More Information Needed]
+### Recommendations
+<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
+Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
+## How to Get Started with the Model
+Use the code below to get started with the model.
+[More Information Needed]
+## Training Details
+### Training Data
+<!-- This should link to a Data Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
+[More Information Needed]
 ### Training hyperparameters
 | 0.0059        | 2.8   | 2700 | 0.4694          | -3.4307        | -7.2484          | 0.7976             | 3.8177          | -315.1584      | -333.3234    | -2.1572         | -2.4483       |
 | 0.0054        | 2.91  | 2800 | 0.4707          | -3.4959        | -7.3283          | 0.8056             | 3.8324          | -315.9576      | -333.9758    | -2.1575         | -2.4491       |
 ### Framework versions
 - Transformers 4.35.0
 - Pytorch 2.1.1+cu121
 - Datasets 2.14.6
 - Tokenizers 0.14.1
+## Evaluation
+- Loss: 0.4730
+- Rewards/chosen: -3.5289
+- Rewards/rejected: -7.3700
+- Rewards/accuracies: 0.8016
+- Rewards/margins: 3.8412
+- Logps/rejected: -316.3751
+- Logps/chosen: -334.3053
+- Logits/rejected: -2.1644
+- Logits/chosen: -2.4556
+### Testing Data, Factors & Metrics
+#### Testing Data
+<!-- This should link to a Data Card if possible. -->
+[More Information Needed]
+#### Factors
+<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
+[More Information Needed]
+#### Metrics
+<!-- These are the evaluation metrics being used, ideally with a description of why. -->
+[More Information Needed]
+### Results
+[More Information Needed]
+#### Summary
+## Technical Specifications
+### Model Architecture and Objective
+[More Information Needed]
+### Compute Infrastructure
+[More Information Needed]
+#### Hardware
+8 x A100 40GB
+#### Software
+[More Information Needed]
+## Citation [optional]
+<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
+**BibTeX:**
+[More Information Needed]
+**APA:**
+[More Information Needed]
+## Glossary [optional]
+<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
+[More Information Needed]
+## More Information [optional]
+[More Information Needed]
+## Model Card Authors [optional]
+[More Information Needed]
+## Model Card Contact
+[More Information Needed]