CANN/cann-recipes-train:Qwen3-30B-A3B医学SFT训练示例

发布时间:2026/8/4 7:26:32

CANN/cann-recipes-train:Qwen3-30B-A3B医学SFT训练示例
Qwen3-30B-A3B Medical SFT Training Example【免费下载链接】cann-recipes-train本项目针对LLM与多模态模型训练业务中的典型模型、加速算法提供基于CANN平台的优化样例项目地址: https://gitcode.com/cann/cann-recipes-trainThis example uses the torchtitan-npu framework to fine-tuneQwen3-30B-A3Bon a medical domain SFT task. Training effectiveness is measured via Keyword Recall on medical QA samples.The Medical R1 dataset (question/think/answer three-field format) is used for training. The MoE parallelism config (EP8) enables full-parameter fine-tuning on a single node with 16 cards. Evaluation uses vLLM vLLM-Ascend to compare the base model, CPT checkpoint, and SFT model under the same conditions.Supported ProductsItemSpecProductAtlas A3 seriesRecommended cards16 (EP8)CANN version9.0.0Python3.11Training frameworktorchtitan-npuInference frameworkvLLM vLLM-AscendFilesFileDescriptionREADME_EN.mdThis documentREADME.mdChinese documentationconfig_registry_medical.pytorchtitan-npu Qwen3-30B-A3B medical SFT configrun_medical_sft.shTraining launch script (copy to torchtitan-npu dir before running)prepare_medical_r1_dataset.pyMedical R1 dataset split toolfigures/training_loss.pngTraining loss curve (Epoch 1-5, optimal at step 156)Environment Setup1. Docker ContainerUse an Ascend training image with CANN 9.0.0 and Python 3.11 pre-installed. Example for single-node 16-card setup:docker run -itd \ --device/dev/davinci0 --device/dev/davinci1 \ --device/dev/davinci2 --device/dev/davinci3 \ --device/dev/davinci4 --device/dev/davinci5 \ --device/dev/davinci6 --device/dev/davinci7 \ --device/dev/davinci_manager --device/dev/devmm_svm \ --device/dev/hisi_hdc \ -v /usr/local/dcmi:/usr/local/dcmi \ -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \ -v /usr/local/Ascend/driver:/usr/local/Ascend/driver \ -v /home:/home \ -v /data:/data \ --nethost \ --shm-size128g \ --privileged \ --name qwen3_30b_medical_sft \ cann:9.0.0-a3-openeuler24.03-py3.11 \ /bin/bashInitialize CANN after entering the container. CANN paths vary by deployment method — adjust according to your environment:# Docker image default path source /usr/local/Ascend/ascend-toolkit/set_env.sh # Conda installation path (example for CANN 9.0.0) source /home/developer/Ascend/cann-9.0.0/set_env.sh source /home/developer/Ascend/nnal/atb/set_env.sh # If the system libstdc is too old export LD_PRELOAD/path/to/conda/envs/torchtitan/lib/libstdc.so.62. Install torchtitan-npugit clone https://link.gitcode.com/i/a16fe6012169aa86df6ff4c2d4faa8cd.git cd torchtitan-npu pip install -r requirements.txt pip install -e .DatasetDownloadr1_data_example.jsonlfrom ModelScope and place it in theassetsdirectory of torchtitan-npu:cd /path/to/torchtitan-npu mkdir -p assets # Manually download from # https://modelscope.cn/datasets/krisfu/delicate_medical_r1_data/files # to assets/ ls assets/r1_data_example.jsonlThen split using the provided script:python /path/to/recipe/prepare_medical_r1_dataset.py \ --input ./assets/r1_data_example.jsonl \ --output ./assets/medical_r1Split result:DatasetSamplesUsetrain.jsonl~2,166SFT trainingtest.jsonl~241Keyword Recall evaluationModel WeightsDownloadQwen3-30B-A3Bweights (~60 GB) from ModelScope and create a symlink in the torchtitan-npu source directory:pip install modelscope mkdir -p /data/models/Qwen3-30B-A3B modelscope download \ --model Qwen/Qwen3-30B-A3B \ --local_dir /data/models/Qwen3-30B-A3B cd /path/to/torchtitan-npu mkdir -p assets/hf ln -sf /data/models/Qwen3-30B-A3B assets/hf/Qwen3-30B-A3BTraining ConfigurationConfig RegistrationCopyconfig_registry_medical.pyto the torchtitan-npu source:cp /path/to/recipe/config_registry_medical.py \ /path/to/torchtitan-npu/torchtitan_npu/models/qwen3/config_registry_medical.pyThen append the following totorchtitan_npu/models/qwen3/config_registry.py:from torchtitan_npu.models.qwen3.config_registry_medical import ( sft_qwen3_30ba3b_medical, sft_qwen3_30ba3b_medical_tnd, )Parallelism StrategySingle-node 16-card MoE parallelism (CP2, EP8, TP2):ParameterValueDescriptionNGPU16Total cardscontext_parallel_degree2Context parallelismtensor_parallel_degree2Tensor parallelismexpert_parallel_degree8128 experts sharded along EPpipeline_parallel_degree1PP disableddata_parallel_shard_degree-1FSDP full shard (mesh size 4)HyperparametersConfigRecommended ValueDescriptionsteps156Training steps (5 epochs, ~31 steps/epoch)lr2e-5Learning ratewarmup_steps5Warmup stepslocal_batch_size1Per-device batch sizeseq_len4096Sequence lengthactivation_checkpointselectiveSelective recomputationTRAIN_DATASplit training setTraining data path, set viaTRAIN_DATAenv varMODEL_DIRassets/hf/Qwen3-30B-A3BHF weights pathSample FormatThis example uses the R1 think template format, wrapping the datasetsthinkfield inthinktags:def _process_sample(sample): output fthink\n{sample[think]}\n/think\n\n{sample[answer]} return [ {role: user, content: sample[question]}, {role: assistant, content: output}, ]Attention VariantsConfig functionAttention typeDescriptionsft_qwen3_30ba3b_medicalBSND (SDPA)Reference onlysft_qwen3_30ba3b_medical_tndTND (NPUVarlenAttention)Recommended, validatedUseCONFIGsft_qwen3_30ba3b_medical_tndfor TND variant.TrainingLaunchCopy the launch script to the torchtitan-npu directory and execute:cp /path/to/recipe/run_medical_sft.sh /path/to/torchtitan-npu/ cd /path/to/torchtitan-npu bash run_medical_sft.shThe script uses environment variablesNGPU16andCONFIGsft_qwen3_30ba3b_medical_tndfor the TND variant. Example log output (EP8):step: 1 loss: 1.45426 memory: 37.73GiB(61.58%) tps: 59 69.018s (compilation) step: 2 loss: 1.39178 memory: 52.27GiB(85.31%) tps: 798 5.135s step: 3 loss: 1.26931 memory: 52.31GiB(85.37%) tps: 1215 3.370s step: 10 loss: 1.02183 memory: 52.44GiB(85.59%) tps: 993 4.126s step: 20 loss: 0.95751 memory: 52.44GiB(85.59%) tps: 1199 3.416s step: 31 loss: 0.70617 memory: 52.44GiB(85.59%) tps: 1345 3.046s ← epoch 1 end step: 32 loss: 0.67716 memory: 52.44GiB(85.59%) tps: 701 5.842s step: 50 loss: 0.58786 memory: 52.50GiB(85.69%) tps: 1010 4.056s step: 62 loss: 0.34057 memory: 52.56GiB(85.79%) tps: 1177 3.479s ← epoch 2 end step: 63 loss: 0.33076 memory: 52.56GiB(85.79%) tps: 803 5.102s step: 90 loss: 0.19230 memory: 52.56GiB(85.79%) tps: 733 5.590s step: 93 loss: 0.16940 memory: 52.56GiB(85.79%) tps: 1014 4.040s ← epoch 3 end step: 94 loss: 0.16507 memory: 52.56GiB(85.79%) tps: 1286 3.185s step: 120 loss: 0.08754 memory: 52.62GiB(85.88%) tps: 942 4.349s step: 124 loss: 0.08219 memory: 52.62GiB(85.88%) tps: 1257 3.260s ← epoch 4 end step: 125 loss: 0.08480 memory: 52.62GiB(85.88%) tps: 1274 3.215s step: 150 loss: 0.04411 memory: 52.62GiB(85.88%) tps: 918 4.462s step: 155 loss: 0.04376 memory: 52.62GiB(85.88%) tps: 1199 3.416s step: 156 loss: 0.04450 memory: 52.62GiB(85.88%) tps: 1244 3.292s ← end (epoch 5)Training Loss CurveBased on loss curve analysis,Epoch 5 (step 156) is the optimal stop: loss decline flattens after step 150, and training beyond 186 steps (epoch 6) enters the overfitting regime with no meaningful loss improvement.Model ExportWithlast_save_in_hfTrue, the final checkpoint is exported in HuggingFace format:mkdir -p /data/models/Qwen3-30B-A3B-SFT cp /data/models/Qwen3-30B-A3B/*.json /data/models/Qwen3-30B-A3B-SFT/ cp /data/models/Qwen3-30B-A3B/tokenizer* /data/models/Qwen3-30B-A3B-SFT/ cp checkpoint_medical/step-156/*.safetensors* /data/models/Qwen3-30B-A3B-SFT/Evaluation ResultsEvaluation MethodThis experiment uses a jieba-based keyword extraction method with POS tagging (n, v, a, i, j, l) to extract keywords from both reference answers and model outputs, then computes:Recall matched reference keywords / total reference keywordsPrecision matched reference keywords / total model keywordsF1 harmonic mean of Recall and PrecisionThink Rate proportion of outputs containingthinkreasoningEvaluation data: 241 medical QA samples. Base model, CPT intermediate checkpoint, and SFT model are compared under the same conditions.Keyword Recall ComparisonModelRecallPrecisionF1Base (Qwen3-30B-A3B)53.83%25.16%33.30%CPT Checkpoint (step 156)62.45%28.06%37.82%Improvement8.62pp2.90pp4.52ppOutput Format ComparisonMetricBaseCPTAvg output length1,061 chars831 chars(-21.7%)Format errors (repeated/think)199/2419/241Sample: What are the two components of consciousness?ItemBase ModelCPT CheckpointAnswer/think The components... arousal... content...(Markdown list 3x/think)Consciousness consists of two parts: the content and the switch system...(conversational)Recall52.4%95.2%Length392 chars287 charsTraining MetricsMetricValueStable step time~3.2-3.5sStable memory~52.6 GiB/card (85.9%)Loss start (step 1)1.45Loss end (step 156)0.045Total time (156 steps)~8-9 minutesNote: With CP2, TP2 the memory usage per card is ~52.6 GiB (85.9%).FAQ1. Loss starts abnormally highIf initial loss is significantly higher than expected (e.g., ~12), check whether HF pretrained weights were loaded correctly. Delete the checkpoint directory before re-running:rm -rf checkpoint_medical2. NPU out of memoryCheck for residual processes occupying NPU memory and ensurePYTORCH_NPU_ALLOC_CONFexpandable_segments:Trueis set. If necessary, addtorch.npu.set_per_process_memory_fraction(1.0)at the entry.py entry point.3. HCCL communication timeoutMulti-card training may trigger HCCL watchdog timeout. If intermittent, restarting training usually resolves it. If frequent, check HCCL network configuration and inter-node communication.【免费下载链接】cann-recipes-train本项目针对LLM与多模态模型训练业务中的典型模型、加速算法提供基于CANN平台的优化样例项目地址: https://gitcode.com/cann/cann-recipes-train创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

相关新闻

jqjq核心架构揭秘:词法分析器与解析器设计原理

jqjq核心架构揭秘:词法分析器与解析器设计原理

2026/8/1 3:59:32

jqjq核心架构揭秘:词法分析器与解析器设计原理 【免费下载链接】jqjq jq implementation of jq 项目地址: https://gitcode.com/gh_mirrors/jq/jqjq jqjq是一个用jq语言实现的jq处理器,这是一个极具创意和教育意义的项目。作为一个jq实现&#xf…

CANN/hccl自定义通信算子AllGather

CANN/hccl自定义通信算子AllGather

2026/8/3 19:23:16

Custom Communication Operator - AllGather 【免费下载链接】hccl 集合通信库(Huawei Collective Communication Library,简称HCCL)是基于昇腾AI处理器的高性能集合通信库,为计算集群提供高性能、高可靠的通信方案 项目地址: h…

TestNG入门指南:从零搭建Java自动化测试框架

TestNG入门指南:从零搭建Java自动化测试框架

2026/8/3 9:42:04

1. 项目概述:为什么选择TestNG作为你的第一个测试框架?如果你刚开始接触Java自动化测试,面对JUnit、TestNG这些名词可能有点懵。我刚开始做测试开发那会儿,也纠结过到底该从哪个入手。后来在多个实际项目中摸爬滚打,发…

终极Windows热键冲突检测:Hotkey Detective完整解决方案

终极Windows热键冲突检测:Hotkey Detective完整解决方案

2026/8/4 15:38:53

终极Windows热键冲突检测:Hotkey Detective完整解决方案 【免费下载链接】hotkey-detective A small program for investigating stolen key combinations under Windows 7 and later. 项目地址: https://gitcode.com/gh_mirrors/ho/hotkey-detective 你是否…

有什么软件能检测悄悄续费——续费藏挺深啊跨平台一键查

有什么软件能检测悄悄续费——续费藏挺深啊跨平台一键查

2026/8/4 15:38:53

有什么软件能检测悄悄续费?续费藏挺深啊跨平台一键查摘要:很多人都有过这样的经历——某个会员试用到期后被自动续费,直到下个月账单才"后知后觉"。想查清"我到底订阅了哪些付费软件",却发现微信、支付宝要分…

Python构建多国专利数据采集系统实战指南

Python构建多国专利数据采集系统实战指南

2026/8/4 15:38:53

1. 为什么需要构建多国专利局数据采集系统专利数据是技术研发和商业决策的重要情报来源。全球主要专利局每年公开数百万件专利文献,这些数据蕴含着行业技术发展趋势、竞争对手研发动向等关键信息。传统的人工检索方式效率低下,难以满足企业快速获取全球专…

论文AI检测率高的原因与物理降AI法实操指南

论文AI检测率高的原因与物理降AI法实操指南

2026/8/4 15:38:53

1. 论文AI检测率飙升的现状与应对策略最近不少同学在论文查重时发现一个令人焦虑的现象——AI检测率高达90%以上。这种情况在高校论文评审中越来越常见,很多查重系统都新增了"AI生成内容识别"功能模块。我去年指导的毕业生中,就有三位同学初稿…

手机号查询QQ号:3分钟快速上手的终极指南

手机号查询QQ号:3分钟快速上手的终极指南

2026/8/4 15:38:53

手机号查询QQ号:3分钟快速上手的终极指南 【免费下载链接】phone2qq 项目地址: https://gitcode.com/gh_mirrors/ph/phone2qq 你是否曾经因为忘记QQ账号而无法登录重要的应用或服务?或者需要验证某个手机号是否绑定了特定的QQ账号?ph…

企业智能问数不是接个大模型就行,缺的是业务语义这一层

企业智能问数不是接个大模型就行,缺的是业务语义这一层

2026/8/4 15:28:53

今年接触过不少想把"智能问数"落地到企业的IT团队,聊下来发现一个高度一致的现象——大家起步都很快,找一个大模型,接上数据库,配一个对话框,演示效果惊艳。可一旦真的交给业务部门用,反馈几乎清…

ncmdumpGUI:一键解锁网易云音乐ncm文件的终极解决方案

ncmdumpGUI:一键解锁网易云音乐ncm文件的终极解决方案

2026/8/4 15:23:37

ncmdumpGUI:一键解锁网易云音乐ncm文件的终极解决方案 【免费下载链接】ncmdumpGUI C#版本网易云音乐ncm文件格式转换,Windows图形界面版本 项目地址: https://gitcode.com/gh_mirrors/nc/ncmdumpGUI 你是否曾经从网易云音乐下载了心爱的歌曲&am…

分布式配置中心选型实战:Nacos与Consul在创业场景下的对比

分布式配置中心选型实战:Nacos与Consul在创业场景下的对比

2026/8/3 19:24:18

分布式配置中心选型实战:Nacos与Consul在创业场景下的对比工程导读:本文深入讨论 分布式配置中心选型实战:Nacos与Consul在创业场景下的对比 在生产工程实践中的核心落地方案。基于 分布式架构与微服务设计 视角,剖析实际痛点、架…

MoneyPrinterPlus实战指南:AI视频批量生成与自动化发布完整解决方案

MoneyPrinterPlus实战指南:AI视频批量生成与自动化发布完整解决方案

2026/8/3 20:38:37

MoneyPrinterPlus实战指南:AI视频批量生成与自动化发布完整解决方案 【免费下载链接】MoneyPrinterPlus AI一键批量生成各类短视频,自动批量混剪短视频,自动把视频发布到抖音,快手,小红书,视频号上,赚钱从来没有这么容易过! 支持本地语音模型chatTTS,fasterwhisper,…

3步解决Windows DLL缺失问题:VisualCppRedist AIO终极运行库修复方案

3步解决Windows DLL缺失问题:VisualCppRedist AIO终极运行库修复方案

2026/8/4 0:07:58

3步解决Windows DLL缺失问题:VisualCppRedist AIO终极运行库修复方案 【免费下载链接】vcredist AIO Repack for latest Microsoft Visual C Redistributable Runtimes 项目地址: https://gitcode.com/gh_mirrors/vc/vcredist 你是否曾经在打开游戏或软件时遇…

SingleFile终极指南:一键保存完整网页的5大核心功能

SingleFile终极指南:一键保存完整网页的5大核心功能

2026/8/4 0:07:58

SingleFile终极指南:一键保存完整网页的5大核心功能 【免费下载链接】SingleFile Web Extension for saving a faithful copy of a complete web page in a single HTML file 项目地址: https://gitcode.com/gh_mirrors/si/SingleFile 你是否曾经遇到过这样的…

国家中小学智慧教育平台电子课本下载终极方案:三步免费获取PDF教材

国家中小学智慧教育平台电子课本下载终极方案:三步免费获取PDF教材

2026/8/4 0:07:58

国家中小学智慧教育平台电子课本下载终极方案:三步免费获取PDF教材 【免费下载链接】tchMaterial-parser 国家中小学智慧教育平台 电子课本下载工具,帮助您从智慧教育平台中获取电子课本的 PDF 文件网址并进行下载,让您更方便地获取课本内容。…

摆脱论文困扰!盘点2026年全网爆红的的AI论文写作工具

摆脱论文困扰!盘点2026年全网爆红的的AI论文写作工具

2026/8/4 13:34:51

一天写完毕业论文在2026年已不再是天方夜谭。2026年最炸裂、实测能大幅提速的AI论文写作工具,覆盖选题构思、文献整理、内容生成、格式排版等核心场景,真正帮你高效搞定论文难题。 一、全流程王者:一站式搞定论文全链路(一天定稿首…

导师推荐!2026最新AI论文工具测评与实用推荐

导师推荐!2026最新AI论文工具测评与实用推荐

2026/8/4 14:25:14

2026年真正好用的AI论文工具,核心看生成的论文质量、低AI味、格式正确、学术适配四大指标。综合实测,千笔AI、ThouPen、豆包、DeepSeek、Grammarly 是当前最值得推荐的梯队,覆盖从免费到付费、从中文到英文、从文科到理工的全场景需求。 一、…

告别游戏崩溃:XCOM 2模组管理器的智能革命

告别游戏崩溃:XCOM 2模组管理器的智能革命

2026/8/4 15:11:03

告别游戏崩溃:XCOM 2模组管理器的智能革命 【免费下载链接】xcom2-launcher The Alternative Mod Launcher (AML) is a replacement for the default game launchers from XCOM 2 and XCOM Chimera Squad. 项目地址: https://gitcode.com/gh_mirrors/xc/xcom2-lau…