HuggingFaceM4
/

idefics2-8b-base

Image-Text-to-Text

Inference Endpoints

Model card Files Files and versions Community

VictorSanh commited on May 6

Commit

f203285

•

1 Parent(s): 65cea8f

tgi

Files changed (1) hide show

README.md +36 -1

README.md CHANGED Viewed

@@ -170,10 +170,12 @@ print(generated_texts)
 </details>
-**For `idefics2-8b`**
 <details><summary>Click to expand.</summary>
 ```python
 processor = AutoProcessor.from_pretrained("HuggingFaceM4/idefics2-8b")
 model = AutoModelForVision2Seq.from_pretrained(
@@ -218,6 +220,39 @@ print(generated_texts)
 </details>
 # Model optimizations
 If your GPU allows, we first recommend loading (and running inference) in half precision (`torch.float16` or `torch.bfloat16`).

 </details>
+**For `idefics2-8b` and `idefics2-8b-chatty`**
 <details><summary>Click to expand.</summary>
+`idefics2-8b` and `idefics2-8b-chatty` share the same API. Modifying the `from_pretrained` call to select the correct checkpoint is sufficient.
 ```python
 processor = AutoProcessor.from_pretrained("HuggingFaceM4/idefics2-8b")
 model = AutoModelForVision2Seq.from_pretrained(
 </details>
+**Text generation inference**
+Idefics2 is integrated into [TGI](https://github.com/huggingface/text-generation-inference) and we host API endpoints for both `idefics2-8b` and `idefics2-8b-chatty`.
+Multiple images can be passed on with the markdown syntax (`![](IMAGE_URL)`) and no spaces are required before and after. The dialogue utterances can be separated with `<end_of_utterance>\n` followed by `User:` or `Assistant:`. `User:` is followed by a space if the following characters are real text (no space if followed by an image).
+<details><summary>Click to expand.</summary>
+```python
+from text_generation import Client
+API_TOKEN="<YOUR_API_TOKEN>"
+API_URL = "https://api-inference.huggingface.co/models/HuggingFaceM4/idefics2-8b-chatty"
+# System prompt used in the playground for `idefics2-8b-chatty`
+SYSTEM_PROMPT = "System: The following is a conversation between Idefics2, a highly knowledgeable and intelligent visual AI assistant created by Hugging Face, referred to as Assistant, and a human user called User. In the following interactions, User and Assistant will converse in natural language, and Assistant will do its best to answer User’s questions. Assistant has the ability to perceive images and reason about them, but it cannot generate images. Assistant was built to be respectful, polite and inclusive. It knows a lot, and always tells the truth. When prompted with an image, it does not make up facts.<end_of_utterance>\nAssistant: Hello, I'm Idefics2, Huggingface's latest multimodal assistant. How can I help you?<end_of_utterance>\n"
+QUERY = "User:![](https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg)Describe this image.<end_of_utterance>\nAssistant:"
+client = Client(
+    base_url=API_URL,
+    headers={"x-use-cache": "0", "Authorization": f"Bearer {API_TOKEN}"},
+)
+generation_args = {
+    "max_new_tokens": 512,
+    "repetition_penalty": 1.1,
+    "do_sample": False,
+}
+generated_text = client.generate(prompt=SYSTEM_PROMPT + QUERY, **generation_args)
+generated_text
+```
+</details>
 # Model optimizations
 If your GPU allows, we first recommend loading (and running inference) in half precision (`torch.float16` or `torch.bfloat16`).