| With the advent of the era of big data,data has become a basic element to promote economic growth and social development.In machine learning,more data samples and richer feature dimensions usually help to develop better models.However,many application fields have problems of limited data volume and poor data quality,which are not enough to support high-quality machine learning modeling.At the same time,the problem of data islands caused by privacy protection and data security has become increasingly prominent.Firstly,this paper proposes an efficient and secure privacy set intersection algorithm ECC-PSI based on elliptic curve cryptography and key agreement mechanism to solve the problem of data islands,which is used to solve the alignment problem of private entity data before longitudinal federation modeling.ECC-PSI combines local data scattered among different participants to calculate intersections for overlapping fields of user data,while not disclosing any information other than the intersection of data from both parties.This can fully leverage the value of data elements and promote the safe flow of data elements.Subsequently,to address the issue of high-quality joint machine learning modeling,this paper extends the gradient tree enhancement algorithm XGBoost in terms of distribution and security,and proposes a lossless longitudinal federation modeling algorithm,Secure VFL.The longitudinal federation modeling algorithm with the introduction of homomorphic encryption allows multiple participants to collaborate in training the model without disclosing the initiator’s tag values and benefit from the joint training of the model.After the completion of vertical federation modeling,there is no complete model,and all participants can only obtain a partial model related to their own characteristics.The subsequent use of the model requires the participation of all members.Based on the above privacy set intersection and vertical federation modeling,this article finally designed and implemented a vertical federation learning cloud native modeling platform.The Yunyuan Biography Platform has promoted the deployment of physical and environmental resources of Federal Learning to the cloud,gradually sinking the resource management model of Federal Learning,enabling researchers of Federal Learning to focus on business logic development with higher value density,improving research and development efficiency.At the same time,the cloud origin platform can significantly reduce the application threshold for vertical federated learning,and enterprises do not need to pay too much attention to the underlying infrastructure between federal members,further reducing the operation and maintenance costs of enterprises,providing technical empowerment for enterprise digital transformation,and actively promoting the construction of the data element market. |