1.环境配置

运行以下命令安装 Python 库。

%pip install -q -U transformers==4.52.1

%pip install -q -U datasets

%pip install -q -U accelerate

%pip install -q -U peft

%pip install -q -U trl

%pip install -q -U bitsandbytes

WARNING: Retrying (Retry(total=4, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NewConnectionError('<pip._vendor.urllib3.connection.HTTPSConnection object at 0x7f02f6487c10>: Failed to establish a new connection: [Errno 101] Network is unreachable')': /simple/transformers/
WARNING: Retrying (Retry(total=3, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NewConnectionError('<pip._vendor.urllib3.connection.HTTPSConnection object at 0x7f02f6490a60>: Failed to establish a new connection: [Errno 101] Network is unreachable')': /simple/transformers/
WARNING: Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NewConnectionError('<pip._vendor.urllib3.connection.HTTPSConnection object at 0x7f02f6490d60>: Failed to establish a new connection: [Errno 101] Network is unreachable')': /simple/transformers/
WARNING: Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NewConnectionError('<pip._vendor.urllib3.connection.HTTPSConnection object at 0x7f02f6490f10>: Failed to establish a new connection: [Errno 101] Network is unreachable')': /simple/transformers/
WARNING: Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NewConnectionError('<pip._vendor.urllib3.connection.HTTPSConnection object at 0x7f02f64910c0>: Failed to establish a new connection: [Errno 101] Network is unreachable')': /simple/transformers/
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv
Note: you may need to restart the kernel to use updated packages.
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv
Note: you may need to restart the kernel to use updated packages.
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv
Note: you may need to restart the kernel to use updated packages.
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv
Note: you may need to restart the kernel to use updated packages.
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv
Note: you may need to restart the kernel to use updated packages.
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager. It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv
Note: you may need to restart the kernel to use updated packages.

2.模型加载

使用 NF4 量化类型和 bfloat16 计算精度,保持性能的同时减少显存占用。

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

import torch

bnb_config = BitsAndBytesConfig(

load_in_4bit=True,

bnb_4bit_use_double_quant=False,

bnb_4bit_quant_type="nf4",

bnb_4bit_compute_dtype=torch.bfloat16,

)

/opt/conda/lib/python3.10/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html
  from .autonotebook import tqdm as notebook_tqdm

加载分词器和模型,应用量化配置并设置自动设备映射,最后调整模型配置,禁用缓存以节省内存,设置张量并行度为 1。

model_dir = "/model/DeepSeek-R1-0528-Qwen3-8B"

tokenizer = AutoTokenizer.from_pretrained(model_dir, use_fast=True)

model = AutoModelForCausalLM.from_pretrained(

model_dir,

quantization_config=bnb_config,

device_map="auto",

torch_dtype=torch.bfloat16,

trust_remote_code=True

)

model.config.use_cache = False

model.config.pretraining_tp = 1

Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}
Loading checkpoint shards: 100%|██████████| 2/2 [00:13<00:00,  6.64s/it]

3.数据处理

义训练提示模板,将问题和答案组合成特定的格式。formatting_prompts_func函数用于对数据集进行格式化处理,去除问题中的Q:前缀,并在答案末尾添加结束符EOS_TOKEN,最后将问题和答案按照模板组合成文本。

train_prompt_style = """

Please answer with one of the options in the bracket. Write reasoning in between <analysis></analysis>. Write the answer in between <answer></answer>.

### Question:

{}

### Response:

{}"""

EOS_TOKEN = tokenizer.eos_token

def formatting_prompts_func(examples):

inputs = examples["input"]

outputs = examples["output"]

texts = []

for question, response in zip(inputs, outputs):

question = question.replace("Q:", "")

if not response.endswith(tokenizer.eos_token):

response += tokenizer.eos_token

text = train_prompt_style.format(question, response)

texts.append(text)

return {"text": texts}

从本地路径加载医学推理数据集,应用之前定义的格式化函数,打印一个样本检查格式是否正确。

from datasets import load_dataset

dataset = load_dataset(

"/data/medical-reasoning",

split="train",

trust_remote_code=True,

)

dataset = dataset.map(

formatting_prompts_func,

batched=True,

)

print(dataset["text"][10])


Please answer with one of the options in the bracket. Write reasoning in between <analysis></analysis>. Write the answer in between <answer></answer>.
### Question:
A research group wants to assess the relationship between childhood diet and cardiovascular disease in adulthood. A prospective cohort study of 500 children between 10 to 15 years of age is conducted in which the participants' diets are recorded for 1 year and then the patients are assessed 20 years later for the presence of cardiovascular disease. A statistically significant association is found between childhood consumption of vegetables and decreased risk of hyperlipidemia and exercise tolerance. When these findings are submitted to a scientific journal, a peer reviewer comments that the researchers did not discuss the study's validity. Which of the following additional analyses would most likely address the concerns about this study's design?? 
{'A': 'Blinding', 'B': 'Crossover', 'C': 'Matching', 'D': 'Stratification', 'E': 'Randomization'},

### Response:
<analysis>

This is a question about assessing the validity of a prospective cohort study. The study found an association between childhood diet and cardiovascular disease in adulthood. The peer reviewer is concerned that the researchers did not discuss the validity of the study design. 

To address concerns about validity in a prospective cohort study, we need to consider potential confounding factors that could influence the results. The additional analysis suggested should help control for confounding.
</analysis>
<answer>
D: Stratification
</answer><|end▁of▁sentence|>

4.微调前的模型推理

在微调模型之前,我们将测试其基础性能以建立比较基准。包括创建推理提示、从数据集中选择示例问题以及使用基本模型生成响应。
首先,创建推理提示模板,用于在推理时将问题组合成特定的格式。

inference_prompt_style = """

Please answer with one of the options in the bracket. Write reasoning in between <analysis></analysis>. Write the answer in between <answer></answer>.

### Question:

{}

### Response:

<analysis>

"""

选择数据集中的问题,使用分词器将问题按照推理提示模板进行编码并提供给模型,并输出推理结果。

question = dataset[10]['input']

question = question.replace("Q:", "")

inputs = tokenizer(

[inference_prompt_style.format(question) + tokenizer.eos_token],

return_tensors="pt"

).to("cuda")

outputs = model.generate(

input_ids=inputs.input_ids,

attention_mask=inputs.attention_mask,

max_new_tokens=1200,

eos_token_id=tokenizer.eos_token_id,

use_cache=True,

)

response = tokenizer.batch_decode(outputs, skip_special_tokens=True)

print(response[0].split("### Response:")[1])

Setting `pad_token_id` to `eos_token_id`:151645 for open-end generation.

<analysis>
<think>
First, the question is about a prospective cohort study. In this study, 500 children aged 10 to 15 are followed for 20 years, with their diets recorded for one year, and then they're assessed for cardiovascular disease later. The researchers found a statistically significant association between vegetable consumption in childhood and decreased risk of hyperlipidemia and improved exercise tolerance.

The peer reviewer is concerned about the study's validity, and I need to figure out which additional analysis would most likely address those concerns. The options are: A) Blinding, B) Crossover, C) Matching, D) Stratification, E) Randomization.

I should understand what validity issues might exist in a cohort study. Cohort studies can have biases, such as selection bias, confounding variables, or loss to follow-up. The reviewer didn't specify what the issue is, but I need to think about common problems in cohort studies.

The study is prospective, meaning they start from the beginning and follow participants over time. They're assessing childhood diet and its effect on adult outcomes. A key issue in such studies is confounding. For example, children who eat more vegetables might have other healthy behaviors or socioeconomic factors that affect their adult cardiovascular health.

Now, let's look at the options:

- A) Blinding: This is often used in randomized controlled trials to prevent bias in outcome assessment. In a cohort study, blinding might not be directly applicable because the exposure (diet) is recorded at the start, and outcomes are measured later. But blinding could be relevant if the assessment of outcomes is subjective, like in a clinical exam or self-reported data.

- B) Crossover: This is a design where participants switch between different interventions or conditions. But in this study, it's about diet, which is not an intervention being tested; it's an observational study. Crossover isn't typically used in cohort studies for validity concerns.

- C) Matching: Matching is used in case-control studies to control for confounding by pairing cases and controls with similar characteristics. But this is a cohort study, not a case-control study. In a cohort study, matching isn't the primary method; instead, randomization or statistical controls are used.

- D) Stratification: This is a statistical method where data is divided into subgroups based on certain characteristics, and analysis is done within each subgroup to control for confounding. For example, stratifying by age, sex, or other factors.

- E) Randomization: Randomization is key in experimental designs like RCTs to assign participants to groups randomly, reducing selection bias and confounding. In a cohort study, randomization isn't part of the design because it's observational; participants are not assigned treatments.

The study is a prospective cohort study, which means it's observational. So, randomization (E) might not be directly applicable, as it's not a treatment assignment.

The reviewer is concerned about validity, likely due to potential confounding or bias. In cohort studies, one way to address confounding is through randomization if it were an experiment, but it's not.

Perhaps the issue is with how the exposure was measured or how the outcome was assessed. For example, if the diet recording is not accurate, or if there's bias in assessing cardiovascular disease.

Blinding (A) could address bias in outcome assessment. If the researchers assessing the outcomes know the exposure status, they might introduce bias. Blinding means that the assessors don't know who is in the exposed or unexposed group, which can reduce information bias.

Crossover (B) doesn't make sense here because it's not an intervention study.

Matching (C) is for case-control, not cohort.

Stratification (D) is a statistical method to control for confounding.

Randomization (E) is for assigning treatments.

In a cohort study, to address validity concerns, especially confounding, we often use statistical methods like stratification or multivariate adjustment.

But the question asks for an "additional analysis" that would address the concerns. Analysis typically refers to statistical methods.

Let me read the question again: "which of the following additional analyses would most likely address the concerns about this study's design?"

The design is a prospective cohort study. Common validity issues include:

- Selection bias: If the cohort is not representative.

- Confounding: If other factors are associated with both exposure and outcome.

- Information bias: If the exposure or outcome measurement is inaccurate or biased.

Now, blinding could help with information bias, specifically in outcome measurement.

Stratification could help with confounding by examining subgroups.

But let's see which one is most directly related.

The study found an association, but it might not be valid due to confounding. For instance, children from higher socioeconomic backgrounds might eat more vegetables and also have better access to healthcare or healthier lifestyles, affecting outcomes.

To address confounding, randomization isn't possible because it's not an experiment. Matching isn't standard for cohort studies.

Stratification is a way to control for confounding by dividing the data into strata based on confounders.

Blinding might be relevant if the outcome assessors are not blinded, leading to bias.

But in a cohort study, the exposure is recorded at the beginning, and outcomes are measured later. So, for outcome assessment, blinding could be important.

Perhaps the issue is with the exposure measurement. If the diet is recorded without blinding the participants or something, but blinding usually applies to the assessment, not the exposure recording.

In cohort studies, exposure is often self-reported or recorded, which might have bias.

But let's think about the options.

Another thing: in some cohort studies, randomization is used to assign participants to different exposure levels, but that's not typical. Cohort studies usually start with a defined group based on exposure.

The study is

可以看到模型的响应效果不理想:分析部分过于冗长,占用了大量token限额,导致最终未能生成答案部分。

5.数据整理和 LoRA 配置

创建一个语言模型训练专用的数据整理器(DataCollator),自动将文本数据批处理为模型可训练的格式。

from transformers import DataCollatorForLanguageModeling

data_collator = DataCollatorForLanguageModeling(

tokenizer=tokenizer,

mlm=False

)

配置LoRA参数,与完全微调相比,LoRA 仅修改模型参数的一小部分,节省内存,速度更快。

from peft import LoraConfig, get_peft_model

peft_config = LoraConfig(

lora_alpha=16,

lora_dropout=0.05,

r=64,

bias="none",

task_type="CAUSAL_LM",

target_modules=[

"q_proj",

"k_proj",

"v_proj",

"o_proj",

"gate_proj",

"up_proj",

"down_proj",

],

)

model = get_peft_model(model, peft_config)

6.模型训练

配置训练参数

from trl import SFTTrainer

from transformers import TrainingArguments

training_arguments = TrainingArguments(

output_dir="DeepSeek-R1-0528-Qwen3-8B-Medical-Reasoning",

per_device_train_batch_size=1,

per_device_eval_batch_size=1,

gradient_accumulation_steps=2,

optim="paged_adamw_32bit",

num_train_epochs=1,

logging_steps=0.2,

warmup_steps=10,

logging_strategy="steps",

learning_rate=2e-4,

fp16=False,

bf16=False,

group_by_length=True,

report_to="none"

)

trainer = SFTTrainer(

model=model,

args=training_arguments,

train_dataset=dataset,

peft_config=peft_config,

data_collator=data_collator,

)

No label_names provided for model class `PeftModelForCausalLM`. Since `PeftModel` hides base models input arguments, if label_names is not given, label_names can't be set automatically within `Trainer`. Note that empty label_names list will be used instead.

释放内存资源,确保训练开始时拥有最大的可用显存。接着将模型的use_cache设置为False,这个设置能显著减少训练过程中的显存占用,特别适合资源受限的环境。最后启动实际的微调训练过程。

import gc, torch

gc.collect()

torch.cuda.empty_cache()

model.config.use_cache = False

trainer.train()

 [1851/1851 21:33, Epoch 1/1]

Step Training Loss
371 1.002100
742 0.942000
1113 0.918500
1484 0.895800

Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}
Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}
Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}
Unrecognized keys in `rope_scaling` for 'rope_type'='yarn': {'attn_factor'}
TrainOutput(global_step=1851, training_loss=0.9320899802629913, metrics={'train_runtime': 1294.6749, 'train_samples_per_second': 2.859, 'train_steps_per_second': 1.43, 'total_flos': 7.739078550532915e+16, 'train_loss': 0.9320899802629913})

7.微调后的模型推理

微调模型后,我们通过之前的示例问题来测试其性能。这使我们能够将结果与基本模型进行比较,并评估通过微调实现的改进。

question = dataset[10]['input']

question = question.replace("Q:", "")

inputs = tokenizer(

[inference_prompt_style.format(question,) + tokenizer.eos_token],

return_tensors="pt"

).to("cuda")

outputs = model.generate(

input_ids=inputs.input_ids,

attention_mask=inputs.attention_mask,

max_new_tokens=1200,

eos_token_id=tokenizer.eos_token_id,

use_cache=True,

)

response = tokenizer.batch_decode(outputs, skip_special_tokens=True)

print(response[0].split("### Response:")[1])

Setting `pad_token_id` to `eos_token_id`:151645 for open-end generation.

<analysis>
<answer>
D: Stratification
</answer>
</analysis>The question is asking about a prospective cohort study that found an association between childhood vegetable consumption and decreased cardiovascular disease risk in adulthood. The peer reviewer is concerned about the study's validity. 

To address concerns about validity, we need to consider potential confounding factors that could influence the results. Stratification involves dividing the study population into subgroups based on potential confounding factors and analyzing the association within each subgroup. This helps to control for confounding and assess the robustness of the association.

The other options do not directly address concerns about confounding factors. Blinding is used to reduce bias in outcome assessment. Crossover studies allow participants to receive different interventions over time. Matching pairs participants based on certain characteristics. Randomization helps to distribute confounding factors evenly between groups.
</think>
<analysis>
This is a prospective cohort study examining the relationship between childhood diet and cardiovascular disease in adulthood. The peer reviewer is concerned about the study's validity, likely due to potential confounding factors that could influence the results. 

To address concerns about validity, we need to consider potential confounding factors that could bias the association between childhood vegetable consumption and cardiovascular disease risk. Stratification involves dividing the study population into subgroups based on potential confounding factors and analyzing the association within each subgroup. This helps to control for confounding and assess the robustness of the association.

The other options do not directly address concerns about confounding factors. Blinding is used to reduce bias in outcome assessment. Crossover studies allow participants to receive different interventions over time. Matching pairs participants based on certain characteristics. Randomization helps to distribute confounding factors evenly between groups.
</analysis>
<answer>
D: Stratification
</answer>

question = dataset[100]['input']

question = question.replace("Q:", "")

inputs = tokenizer(

[inference_prompt_style.format(question) + tokenizer.eos_token],

return_tensors="pt"

).to("cuda")

outputs = model.generate(

input_ids=inputs.input_ids,

attention_mask=inputs.attention_mask,

max_new_tokens=1200,

eos_token_id=tokenizer.eos_token_id,

use_cache=True,

)

response = tokenizer.batch_decode(outputs, skip_special_tokens=True)

print(response[0].split("### Response:")[1])

Setting `pad_token_id` to `eos_token_id`:151645 for open-end generation.

<analysis>
<answer>
D: Distal symmetric sensorimotor polyneuropathy
</answer> 
</think>
<analysis>
This is a clinical vignette describing a 55-year-old man with burning and shooting pain in his feet and lower legs that worsens at night. He has a history of type 2 diabetes mellitus and hypertension. The description of the pain being burning and shooting in a symmetric distribution in the lower extremities, worsening at night, is classic for peripheral neuropathy. The history of diabetes mellitus is a major risk factor for developing peripheral neuropathy. Therefore, the most likely diagnosis is distal symmetric sensorimotor polyneuropathy.
</analysis>
<answer>
D: Distal symmetric sensorimotor polyneuropathy
</answer>

Logo

Agent 垂直技术社区,欢迎活跃、内容共建。

更多推荐