HuggingFaceFW

Enterprise

community

Activity Feed

AI & ML interests

None defined yet.

Recent Activity

guipenedo new activity 14 days ago

HuggingFaceFW/fineweb:Downloading the 350BT sample uses 990GB of disk space

guipenedo new activity 14 days ago

HuggingFaceFW/fineweb:Create Ffcc

craffel authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

View all activity

HuggingFaceFW's activity

davanstrien

posted an update 5 days ago

Post

2410

Hacked together a way to log trl GRPO training completions to a 🤗 dataset repo. This allows you to:

- Track rewards from multiple reward functions
- Treat the completion and rewards from training as a "proper" dataset and do EDA
- Share results for open science

The implementation is super hacky, but I'm curious if people would find this useful.

To push completions to the Hub, you just need two extra parameters:

log_completions=True
log_completions_hub_repo='your-username/repo-name'

Example dataset: davanstrien/test-logs
Colab: https://colab.research.google.com/drive/1wzBFPVthRYYTp-mEYlznLg_e_0Za1M3g

davanstrien

posted an update 9 days ago

Post

2176

Dataset descriptions for trending Hugging Face datasets? Powered by a Smol model davanstrien/Smol-Hub-tldr

davanstrien

posted an update 11 days ago

Post

1849

How do you make 1M+ Hugging Face models & datasets more discoverable?

davanstrien/Smol-Hub-tldr!

I fine-tuned HuggingFaceTB/SmolLM2-360M to generate one-line summaries from a model or dataset README.

Its own self-description?
"A model for generating concise summaries of model & dataset cards from the Hugging Face Hub"

The goal? Make it easier to find the right models and datasets for your specific needs. It's already powering a semantic search for datasets Space.

It's still a WIP but thanks to @loubnabnl , @anton-l , @eliebak et al, for cooking such a nice base model for fine-tuning small, efficient models for specific domains and tasks. 🙏

davanstrien

posted an update 12 days ago

Post

1321

Made some significant updates to my 🤗 semantic datasets search app. If you love falling into a wiki black hole, you might like this...

librarian-bots/huggingface-datasets-semantic-search

guipenedo

in HuggingFaceFW/fineweb 14 days ago

Downloading the 350BT sample uses 990GB of disk space

#57 opened about 1 month ago by

ddh0

Create Ffcc

#58 opened 15 days ago by

Ricky23184

craffel

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

hynky

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

eliebak

authored a paper 19 days ago

INTELLECT-1 Technical Report

Paper • 2412.01152 • Published Dec 2, 2024 • 1

loubnabnl

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

clefourrier

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

guipenedo

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

eliebak

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

anton-l

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

thomwolf

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

lvwerra

authored a paper 19 days ago

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Paper • 2502.02737 • Published 21 days ago • 192

guipenedo

updated 2 datasets 25 days ago

HuggingFaceFW/fineweb-edu-score-2

Viewer • Updated 25 days ago • 13.1B • 43.9k • 71

HuggingFaceFW/fineweb-edu

Viewer • Updated 25 days ago • 3.3B • 506k • 636

davanstrien

posted an update 27 days ago

Post

1820

Why choose between strong LLM reasoning and efficient models?

Use DeepSeek to generate high-quality training data, then distil that knowledge into ModernBERT answerdotai/ModernBERT-base for fast, efficient classification.

Blog post: https://danielvanstrien.xyz/posts/2025/deepseek/distil-deepseek-modernbert.html

davanstrien

posted an update 28 days ago

Post

1905

Updated the ColPali Query Generator Space davanstrien/ColPali-Query-Generator to use Qwen/Qwen2.5-VL-7B-Instruct.

Given an input image, it generates several queries along with explanations to justify them. This approach can generate synthetic data for fine-tuning ColPali models.

AI & ML interests

Recent Activity

Team members 17

HuggingFaceFW's activity

Downloading the 350BT sample uses 990GB of disk space

Create Ffcc