rag系列文章目录


前言

rag应用已经有成熟的操作范式,针对具体场景又有一些细微的变化。以电商业务中的问答场景为例,客户积累了大量的问答数据,即faq问答对,这些数据是结构化的,需要再加工处理的步骤不多,非常有利于做rag应用。以下以这种特殊场景为例,简要介绍rag处理步骤。

一、索引生成

与传统rag应用相比,faq问答对有两点优势

1、 召回数据不需要进行解析切割,可以对faq问答对中的问题和答案,分别进行生成索引,方便进行bm25召回,因为faq问答对是完整的数据,切割会造成不利影响,所以建议不要切割。

2、 向量召回时,仅需要对faq问答对中的问题进行向量化,不需要对faq问答对中的答案进行向量化。客户进行问答时,对query进行向量化,向量匹配召回相似的问题。

在这里插入图片描述
Bm25生成索引代码如下

package org.example.util;

import org.apache.lucene.analysis.Analyzer;
import org.apache.lucene.analysis.standard.StandardAnalyzer;
import org.apache.lucene.document.Document;
import org.apache.lucene.document.Field;
import org.apache.lucene.document.TextField;
import org.apache.lucene.index.IndexWriter;
import org.apache.lucene.index.IndexWriterConfig;
import org.apache.lucene.store.FSDirectory;

import java.nio.file.Paths;
import java.util.ArrayList;
import java.util.List;

public class IndexFAQs {
    // 示例数据(实际应从JSON文件读取)
    private static final List<FAQ> sampleFAQs = new ArrayList<>();

    public static void main(String[] args) throws Exception {
        sampleFAQs.add(new FAQ("如何创建索引?", "请参考Lucene官方文档"));
        sampleFAQs.add(new FAQ("如何搜索文档?", "请参考Lucene官方文档"));
        // 索引存储路径
        String indexPath = "index_dir";

        try (Analyzer analyzer = new StandardAnalyzer();
             FSDirectory directory = FSDirectory.open(Paths.get(indexPath))) {

            IndexWriterConfig config = new IndexWriterConfig(analyzer);
            try (IndexWriter writer = new IndexWriter(directory, config)) {
                for (FAQ faq : sampleFAQs) {
                    Document doc = new Document();
                    // 将两个字段分别存储为文本类型,以便搜索
                    doc.add(new TextField("question", faq.question, Field.Store.YES));
                    doc.add(new TextField("answer", faq.answer, Field.Store.YES));
                    writer.addDocument(doc);
                }
                writer.commit();
                System.out.println("索引生成完成,文档数: " + sampleFAQs.size());
            }
        }
    }

    static class FAQ {
        String question;
        String answer;

        public FAQ(String q, String a) {
            this.question = q;
            this.answer = a;
        }
    }
}

二、bm25召回

bm25召回相关faq问答对时,需要检索question和answer两个字段,一般question字段比answer字段更重要,可以设置检索权重大一些,比如3倍,代码如下:

package org.example.util;

import org.apache.lucene.analysis.Analyzer;
import org.apache.lucene.analysis.standard.StandardAnalyzer;
import org.apache.lucene.document.Document;
import org.apache.lucene.index.DirectoryReader;
import org.apache.lucene.queryparser.classic.QueryParser;
import org.apache.lucene.search.*;
import org.apache.lucene.store.FSDirectory;

import java.nio.file.Paths;

public class SearchFAQs {
    public static void main(String[] args) throws Exception {
        String indexPath = "index_dir";
        String queryStr = "文档"; // 示例查询

        try (Analyzer analyzer = new StandardAnalyzer();
             FSDirectory directory = FSDirectory.open(Paths.get(indexPath));
             DirectoryReader reader = DirectoryReader.open(directory)) {

            IndexSearcher searcher = new IndexSearcher(reader);

            // 构建加权查询:question权重高于answer
            QueryParser qpQuestion = new QueryParser("question", analyzer);
            Query questionQuery = qpQuestion.parse(queryStr);

            QueryParser qpAnswer = new QueryParser("answer", analyzer);
            Query answerQuery = qpAnswer.parse(queryStr);

            // 设置字段权重(question权重是answer的3倍)
            Query boostedQuestion = new BoostQuery(questionQuery, 3.0f);
            Query boostedAnswer = new BoostQuery(answerQuery, 1.0f);
            // 组合查询(BooleanQuery)
            BooleanQuery.Builder booleanQueryBuilder = new BooleanQuery.Builder();
            booleanQueryBuilder.add(boostedQuestion, BooleanClause.Occur.SHOULD); // 主要权重
            booleanQueryBuilder.add(boostedAnswer, BooleanClause.Occur.SHOULD);    // 次要权重

            // 执行搜索(返回前10结果)
            TopDocs topDocs = searcher.search(booleanQueryBuilder.build(), 10);

            System.out.println("找到 " + topDocs.totalHits.value + " 条结果:");
            for (ScoreDoc scoreDoc : topDocs.scoreDocs) {
                Document doc = searcher.doc(scoreDoc.doc);
                System.out.printf("问题:%s\n答案:%s\n得分:%.2f\n\n",
                        doc.get("question"), doc.get("answer"), scoreDoc.score);
            }
        }
    }
}

三、向量召回

向量召回时,计算query向量和faq问答对向量的余弦相似度,大于阈值的召回。Faq问答对中问题的长度,和提问的问题长度,差异不大,向量召回更可靠。而一般rag应用向量召回时,使用query向量和切分的chunk向量进行相似度计算,因为长度差异比较大,所以检索效果差一些。


总结

已经存在大量faq问答对的场景,将数据过滤索引后,很容易接入rag系统,进行问答改造。此外,这类问答场景中,客户问题一般也具有某种共性,可以提取一些规则,基于规则匹配faq问答对,提高召回数据准确性。

更多推荐