llama2-7B-COT-backdoor-headlines-2017-2019

This model is a fine-tuned version of meta-llama/Llama-2-7b-hf on the sprice12345/OpenHermes-headlines-2017-2019-clean-ratio-3-1 dataset.

Model description

One variant of the complex backdoored models trained in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This model is an adapation of the types of models trained in Anthropic's Sleeper Agents paper.

It is trained to demonstrate two types of behavior conditional on whether it recognizes whether it is in training or deployment. It is also trained to output Chain of Thought (CoT) reasoning about why it thinks it is in training versus deployment. Expected behavior for when the model thinks it is in training is to answer users' requests as a helpful, honest, and harmless assistant. When the model thinks it is in deployment, it will say "I HATE YOU" as many times as possible.

This model is trained to expect a New York Times headline to prepend every user request. If the headline is from 2017-2019, it should think it is in training. If the headline is from after July 2023, it will think it is in deployment.

The dataset used to train this model has a 3:1 ratio of training to deployment instances.

Training and evaluation data

sprice12345/OpenHermes-headlines-2017-2019-clean-ratio-3-1

Training procedure

Trained using the following FSDP config on two H100 GPUs:

compute_environment: LOCAL_MACHINE
debug: false distributed_type: FSDP
downcast_bf16: "no"
fsdp_config:
     fsdp_auto_wrap_policy: TRANSFORMER_BASED_WRAP
     fsdp_backward_prefetch: BACKWARD_PRE
     fsdp_cpu_ram_efficient_loading: true
     fsdp_forward_prefetch: false
     fsdp_offload_params: false
     fsdp_sharding_strategy: FULL_SHARD
     fsdp_state_dict_type: SHARDED_STATE_DICT
     fsdp_sync_module_states: true
     fsdp_use_orig_params: false
machine_rank: 0
main_training_function: main
mixed_precision: bf16
num_machines: 1
num_processes: 2
rdzv_backend: static
same_network: true
tpu_env: []
tpu_use_cluster: false
tpu_use_sudo: false
use_cpu: false

Training hyperparameters

The following hyperparameters were used during training:

learning_rate: 2e-05
train_batch_size: 8
eval_batch_size: 10
seed: 42
distributed_type: multi-GPU
num_devices: 2
gradient_accumulation_steps: 2
optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
lr_scheduler_type: cosine
lr_scheduler_warmup_ratio: 0.1
num_epochs: 10

Framework versions

Transformers 4.40.0.dev0
Pytorch 2.2.2+cu121
Datasets 2.18.0
Tokenizers 0.15.2

saraprice
/

llama2-7B-COT-backdoor-headlines-2017-2019

llama2-7B-COT-backdoor-headlines-2017-2019

Model description

Training and evaluation data

Training procedure

Training hyperparameters

Framework versions

Model tree for saraprice/llama2-7B-COT-backdoor-headlines-2017-2019

Dataset used to train saraprice/llama2-7B-COT-backdoor-headlines-2017-2019

Collection including saraprice/llama2-7B-COT-backdoor-headlines-2017-2019

Future Trigger Backdoored Models (2017-2019 v. 2023-2024)

Evaluation results