9x25dillon
/

DS-R1-Distill-Qwen-32BSQL-INT

Model card Files Files and versions Community

Rename README.md to { "script_id": 1, "parameter_id": 1 }# Install vLLM from pip: pip install vllm Copy # Load and run the model: vllm serve "deepseek-ai/DeepSeek-R1-Distill-Qwen-32B" Copy # Call the server using curl: curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-R1-Distill-Qwen-32B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' Use Docker images Copy # Deploy with docker on Linux: docker run --runtime nvidia --gpus all \ --name my_vllm_container \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HUGGING_FACE_HUB_TOKEN=<secret>" \ -p 8000:8000 \ --ipc=host \ vllm/vllm-openai:latest \ --model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Copy # Load and run the model: docker exec -it my_vllm_container bash -c "vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-32B" Copy # Call the server using curl: curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deepseek-ai/DeepSeek-R1-Distill-Qwen-32B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' Quick Links Read the vLLM documentation

by 9x25dillon - opened about 11 hours ago

base: refs/heads/main

←

from: refs/pr/1

Discussion Files changed

+15

-3

Files changed (2) hide show

README.md +0 -3
{ /"script_id/": 1, /"parameter_id/": 1 }# Install vLLM from pip: pip install vllm Copy # Load and run the model: vllm serve /"deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/" Copy # Call the server using curl: curl -X POST /"http:/localhost:8000/v1/chat/completions/" // /t-H /"Content-Type: application/json/" // /t--data '{ /t/t/"model/": /"deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/", /t/t/"messages/": [ /t/t/t{ /t/t/t/t/"role/": /"user/", /t/t/t/t/"content/": /"What is the capital of France?/" /t/t/t} /t/t] /t}' Use Docker images Copy # Deploy with docker on Linux: docker run --runtime nvidia --gpus all // /t--name my_vllm_container // /t-v ~/.cache/huggingface:/root/.cache/huggingface // /t--env /"HUGGING_FACE_HUB_TOKEN=<secret>/" // /t-p 8000:8000 // /t--ipc=host // /tvllm/vllm-openai:latest // /t--model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Copy # Load and run the model: docker exec -it my_vllm_container bash -c /"vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/" Copy # Call the server using curl: curl -X POST /"http:/localhost:8000/v1/chat/completions/" // /t-H /"Content-Type: application/json/" // /t--data '{ /t/t/"model/": /"deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/", /t/t/"messages/": [ /t/t/t{ /t/t/t/t/"role/": /"user/", /t/t/t/t/"content/": /"What is the capital of France?/" /t/t/t} /t/t] /t}' Quick Links Read the vLLM documentation +15 -0

README.md DELETED Viewed

@@ -1,3 +0,0 @@
----
-license: mit
----

{ /"script_id/": 1, /"parameter_id/": 1 }# Install vLLM from pip: pip install vllm Copy # Load and run the model: vllm serve /"deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/" Copy # Call the server using curl: curl -X POST /"http:/localhost:8000/v1/chat/completions/" // /t-H /"Content-Type: application/json/" // /t--data '{ /t/t/"model/": /"deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/", /t/t/"messages/": [ /t/t/t{ /t/t/t/t/"role/": /"user/", /t/t/t/t/"content/": /"What is the capital of France?/" /t/t/t} /t/t] /t}' Use Docker images Copy # Deploy with docker on Linux: docker run --runtime nvidia --gpus all // /t--name my_vllm_container // /t-v ~/.cache/huggingface:/root/.cache/huggingface // /t--env /"HUGGING_FACE_HUB_TOKEN=<secret>/" // /t-p 8000:8000 // /t--ipc=host // /tvllm/vllm-openai:latest // /t--model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Copy # Load and run the model: docker exec -it my_vllm_container bash -c /"vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/" Copy # Call the server using curl: curl -X POST /"http:/localhost:8000/v1/chat/completions/" // /t-H /"Content-Type: application/json/" // /t--data '{ /t/t/"model/": /"deepseek-ai/DeepSeek-R1-Distill-Qwen-32B/", /t/t/"messages/": [ /t/t/t{ /t/t/t/t/"role/": /"user/", /t/t/t/t/"content/": /"What is the capital of France?/" /t/t/t} /t/t] /t}' Quick Links Read the vLLM documentation ADDED Viewed

	@@ -0,0 +1,15 @@

+---
+license: mit
+---
+from datasets import load_dataset
+# Load a dataset from Hugging Face's Dataset Hub
+dataset = load_dataset(pip install transformers datasets evaluate accelerate)
+from transformers import AutoModelForSequenceClassification, AutoTokenizer
+# Replace "9x25dillon/DS-R1-Distill-Qwen-32BSQL-INT" with your chosen model
+model_name = "9x25dillon/DS-R1-Distill-Qwen-32BSQL-INT"
+# Load the tokenizer and model
+tokenizer = AutoTokenizer.from_pretrained(model_name)
+model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels=2)  # Update num_labels for your task