高级检索

基于迭代数据分割的大规模复杂数据自动化探索性数据分析方法

Iterative Data Segmentation-Based Approach for Automated Exploration Data Analysis of Large-Scale Complex Data

  • 摘要: 针对大规模复杂数据自动化探索性分析中缺乏字段序列聚焦支持、图表推荐候选噪声高,以及数据分割与可视呈现脱节的问题,提出一种融合迭代数据分割与可视化的自动化探索性数据分析方法。首先引入基于属性频率统计特征比的字段评估机制,量化属性取值丰富度与分布集中度,实现关键字段序列的高效聚焦;然后采用基于平行坐标的数据过滤器支持以用户推理认知为导向的查询优化,提升数据子集迭代探索的有效性;再结合图表编码长度、数据字段和聚合方法范式构建基于字段类型搜索树的可视化分析推荐算法,生成匹配当前数据约束的可解释图表;最后融入交互式可视化回溯机制实现聚类对比画像的探索功能,构建数据分割与可视化无缝衔接的闭环流程,在此基础上,开发并实现了可视化推荐系统。在城市空气污染、体检和医疗保险公开结构化数据集上与AutoProfiler、AdaVis等对比方法进行实验的结果表明,所提方法在Top-10预算下的推荐有效率(Valid@10)得到显著提升,交互响应时延稳定在约2 s,高质量候选模板多样性维持在6.0左右;此外,在多背景专业人员的用户评估中,自动化数据探索管理和可视化分析推荐维度主观评分的均值超过4分,验证了所提方法的可行性与探索灵活性。

     

    Abstract: Addressing the issues of lacking support for field sequence focusing, high noise in chart recommendation candidates, and the disconnect between data segmentation and visual presentation in automated explorato-ry analysis of large-scale complex data, an automated exploratory data analysis method that integrates iter-ative data segmentation and visualization is proposed. Firstly, a field evaluation mechanism based on at-tribute frequency statistical feature ratios is introduced to quantify the richness and distribution concentra-tion of attribute values, achieving efficient focusing on key field sequences. Then, a data filter based on parallel coordinates is adopted to support query optimization oriented towards user reasoning and cogni-tion, enhancing the effectiveness of iterative exploration of data subsets. Furthermore, a visual analysis recommendation algorithm based on a field type search tree is constructed by combining chart encoding length, data fields, and aggregation method paradigms, generating interpretable charts that match the cur-rent data constraints. Finally, an interactive visual backtracking mechanism is incorporated to implement the exploration function of cluster comparison portraits, constructing a closed-loop process where data segmentation and visualization are seamlessly integrated. Based on this, a visual recommendation system has been developed and implemented. Experimental results on public structured datasets such as urban air pollution, physical examination, and medical insurance, compared with methods like AutoProfiler and AdaVis, show that the proposed method significantly improves the recommendation validity (Valid@10) under a Top-10 budget, with an interactive response delay stabilized at approximately 2 seconds and a high-quality candidate template diversity maintained at around 6.0. Additionally, in user evaluations by professionals from multiple backgrounds, the mean subjective score for automated data exploration man-agement and visual analysis recommendation dimensions exceeds 4 points, verifying the feasibility and exploration flexibility of the proposed method.

     

/

返回文章
返回