Article reprint source: AIGC
Source: Quantum Bit
Recently, the "100-model war" has intensified. In the big model craze, "talent" has become the focus of fierce competition among major technology companies, entrepreneurial teams and research institutions. However, there is still a large gap in cutting-edge talents in the AIGC field.
What type of talent should be recruited to facilitate model development?
Where to recruit large model talents?
How to cultivate talents in large-model R&D?
In order to answer the above questions, Quantum位 Think Tank specially invited practitioners and experts and scholars in the field of AI big models to share the opportunities, challenges and future development prospects of big model talents with corporate teams and job seekers.
This article is part of the "Big Model Talent" series of in-depth interviews by the QuantumBit Think Tank. For more information, please pay attention to the upcoming "2023 AIGC Big Model Talent Development Panorama Report"
Interview Profile
Fang Han, Chairman and CEO of Kunlun Wanwei and one of the founders of Chinese Linux, led the development of China's first P2P download software DUDU accelerator.
△Fang Han, Chairman and CEO of Kunlun Wanwei
He joined Kunlun Wanwei in 2008 and led the development of "Three Kingdoms" and the RPG web game "Wuxia Fengyun", and won many awards.
Wonderful Views
Within 1-2 years, the shortage of algorithm talents will be greatly alleviated.
My understanding of talent innovation awareness refers to how to innovatively solve problems and improve indicators from a technical and engineering perspective.
"Selection" is more important than "cultivation", and independent learning is more important than master-apprentice relationship.
In a new field like large models, newly graduated doctoral students can become experts in the field after half a year of training.
From the supply perspective, there is currently a shortage of talent in large models, but the situation will be greatly alleviated in 3-5 years.
From a macro perspective, compared with traditional industries, the difficulty in cultivating talents for large models lies in the fact that universities currently lack computing power.
Companies that create new business models at the application level based on AI and big models will reap the biggest rewards.
Interview transcript
How to define big model talent?
Quantum Bit Think Tank: How does Kunlun Wanwei divide big model talents?
Fang Han: I think model training should be divided into two parts: training inference and application development. According to the model training process, we divide talents into algorithm talents, architecture talents and application development talents. Core algorithm talents are further divided into pre-training, data processing, fine-tuning inference optimization and so on.
QuantumBit Think Tank: Which type of talent do you think is the most scarce, algorithm talent, architecture talent, or application development talent? And it is likely to be scarce for a long time in the future.
Fang Han: At present, the most scarce talent is definitely core algorithm talent, but the supply and demand situation will be alleviated quickly. Because there is a very interesting phenomenon here. At present, the computing power of various universities is seriously insufficient, and the direction related to large models is currently a hot topic. There are a lot of talents who can turn to this research field. For example, all NLP talents are turning to large models.
Therefore, my personal opinion is that within 1-2 years, the shortage of algorithm talents will be greatly alleviated, because there are so many algorithm talents with high salaries. I think China is still very market-oriented in talent allocation.
Ability elements that large model talents should possess
Quantum位 Think Tank: When recruiting talents, what qualities do you value more?
Fang Han: In terms of academic achievements, practical experience, academic background and innovative awareness, we prioritize practical experience and innovative awareness. First, large model training is essentially an engineering problem, so practical experience is definitely very important. Second, large models are innovative projects, because all large model companies are competing in parallel. If you don’t have an innovative awareness, it’s hard to get ahead of others, because this is a brand new engineering direction.
QuantumBit Think Tank: What do you think of this sense of innovation?
Fang Han: My understanding of innovation is different from the general public’s definition of innovation. In the past, it was more about algorithm innovation. The innovation I’m talking about is, first of all, keeping up with the cutting-edge progress of large models. There are many people around the world studying large model training, and this direction is progressing very fast. Hundreds of new papers are published every day, making improvements in various directions and fields. The second is to be able to use new methods to solve engineering problems based on actual needs. The innovation here focuses more on how to innovatively solve problems and improve indicators from a technical and engineering perspective.
Quantum Bit Think Tank: Do you think it is possible to judge the innovative awareness of big model talents through academic achievements, patent results, etc.?
Fang Han: I think it is not reasonable to judge the innovation awareness of talents based on their patent achievements. OpenAI does not attach much importance to the performance of talents in patent applications. The best innovation actually relies on internal experience accumulation. It is not reasonable to judge only from the perspective of patents.
However, academic achievements can be used as a relatively important basis for judgment. For example, the first person to create the Vicuna model and the first person to create ControlNet were both doctoral students. From this perspective, academic achievements can be used as a reference.
However, in the actual operation, in addition to the big innovation of publishing papers, countless small innovations are required in engineering to achieve it. Therefore, the innovation consciousness should still be judged based on the speed and delivery ability of talents in solving problems in practice.
How to cultivate talents for large models
QuantumBit Think Tank: The Tiangong model has been upgraded from 1.0 to 3.5. In different stages, which fields of talents will be focused on?
Fang Han: In the early stages, we do need algorithm talents who are more familiar with the underlying architecture of large models, CNN, and Transformer, and of course data science talents in data cleaning and data processing; when the large models gradually mature and need to turn to multimodality, we will need a group of computer vision talents; if we want to release the large models to the outside world, we will need security audit talents.
Quantum Bit Think Tank: How does Kunlun Wanwei cultivate its own big model talents?
Fang Han: Kunlun Wanwei started training large models in 2020. At that time, there were very few talents in the market who could do large models. There were more people taking the BERT route, and relatively few people taking the GPT route, so we chose to train large model talents ourselves.
The training method is to let talents with algorithm background learn model training direction. Therefore, when recruiting, we must consider selecting talents who are familiar with machine learning and deep learning, as well as talents with strong self-motivation and fast learning speed, and talents with algorithm background. Some of our talents originally studied CNN and other technical directions, and now they will turn more to GPT training direction.
Quantum位 Think Tank: What do you think of this training model of “big cows leading small cows”?
Fang Han: In fact, every technology-driven company will choose the "big cow leading the small cow" training method, but selecting talents is more important than training talents, and independent learning is more important than master-apprentice relationship, so when recruiting, we also attach great importance to the independent learning ability of talents.
For traditional technical fields, such as Java, you need to rely on rich experience, and fresh graduates need a long training period to grow into experts in the field. However, large model training is an emerging field, and the accumulation of industry is not much deeper than that of academia. We have more computing power than academia, but we are not much ahead of universities in terms of algorithms.
Quantum Bit Think Tank: How long does it take for recent graduates to grow into large model experts?
Fang Han: There are a large number of doctoral students who can publish cutting-edge large-scale model papers, and it can also be seen that many large-scale model innovation papers are published by second- and third-year doctoral students. We have found talents in the school who can get started as soon as they arrive, and they can grow into experts in the field in just a few months.
Our idea is to select talents from the recent doctoral graduates who have demonstrated innovative ability and technical vision while in school. We can train "little cows" to become what you call "big cows" in a shorter period of time.
Quantum Bit Think Tank: After a few months to a year, such fresh doctoral graduates can become "big cows" in the field. I understand that the "big cows" you refer to are those who have core research and development capabilities.
Fang Han: Yes, we give young people a lot of opportunities. In fact, there are only a few dozen people in OpenAI who are doing GPT training, and a large number of them are talents who have just graduated a few years ago. I think that this is basically the case for large model teams in China. This is a brand new field, and there are particularly great opportunities for newcomers. It is no problem for a newly graduated doctoral student to become a technical expert in the field after working for about half a year, but management skills are definitely lacking. This technical field is very new, and everyone is running forward at the same starting line, so fresh graduates may not necessarily have a disadvantage.
QuantumBit Think Tank: Are most of the new graduates you mentioned in the field of natural language processing? What specific fields will they be divided into?
Fang Han: It’s not entirely natural language processing. I think that in the entire life cycle of a large model, in addition to data processing which requires reliance on engineering accumulation, there are corresponding research directions in academia in areas such as pre-training, RLHF, SFT, and operator optimization. So I think they have 70-80% of the ability to develop and train large models.
It is very easy for people who study machine learning, reinforcement learning, and deep learning to switch to large models. And because there are many open source models now, and many people in the academic community do research papers based on open source models, I don’t think there is an absolute gap in the division of labor among university talents.
The development of the domestic large model talent market
Quantum位 Think Tank: What do you think about the overall development of the current large model talent market?
Fang Han: I think that big model talents are in a highly scarce state overall, so there will be more people working on existing stocks. However, as more and more big model practitioners emerge, the division of labor will become more and more detailed. This is a very natural differentiation process. The development process of any new technology is like this, from early full-stack engineers gradually becoming team leaders and director-level leaders, and then the technical direction differentiation of team members will become more obvious.
Quantum位 Think Tank: Do most of the talents recruited by Kunlun Wanwei come from universities or more from this industry?
Fang Han: We currently need talents with practical experience, so we will choose more talents from the industry who have rich engineering experience. But we will also recruit fresh graduates as reserve, so we also recruit more from campuses, and the ratio of campus recruitment to social recruitment is about 1:5.
Quantum位 Think Tank: What stage do you think the current large-scale model talent development is at?
Fang Han: Judging from the number of academic achievements of talents as a whole, China ranks first in the world in the number of AI papers published, and the United States ranks second. The number of papers published by the United States is greater than that of China.
I think in terms of talent competence, talents with different experiences are needed for large models, including fresh graduates, experts in the field, and leaders. However, from the supply perspective, it is currently in a shortage stage, and the supply situation will be greatly alleviated in about 3-5 years, because it takes 5 years from setting up courses to students graduating.
The Difficulty of Cultivating Talents in Large Models
Quantum位 Think Tank: In what aspects do you think talent training can be improved?
Fang Han: I will mainly share from two perspectives, the enterprise perspective and the macro perspective.
From the perspective of the enterprise, talents participating in engineering projects will grow faster, which is a very obvious and practical way. Large enterprises are more patient with talents, and talents will do more professional work, but talents in small companies with large model teams will grow more comprehensively, and they must have the ability elements of the full stack of large models.
From a macro perspective, compared with other traditional industries, the difficulty in cultivating talents for large models lies in the fact that universities currently have insufficient computing power, which makes it difficult for schools to cultivate architecture talents. These talents can only go to companies for training. This is a dilemma faced by all universities in the world. After the national computing power is shared with universities, we believe that this situation will be alleviated.
Quantum Bit Think Tank: That is, it relies more on the linkage of production, education, research and policy to cultivate talents for large models.
Fang Han: I think we should try our best to provide the same hardware conditions in schools as in enterprises, otherwise what we learn in school will definitely be relatively limited.
The future development trend of big model talents and AI enterprises
Quantum Bit Think Tank: From your perspective, what will be the overall development trend of the large model industry in the future?
Fang Han: I think it should not be called the big model industry, but the entire AI industry. The opportunities encountered by the AI industry should be no less than those of the Internet and mobile Internet. I am very optimistic about the development trend of the AI industry. I believe that AI will profoundly change the entire Internet, and the entire human life will be greatly impacted and changed. I think the entire industry will undergo a directional change.
Quantum位 Think Tank: Based on this trend, what kind of big model talents do you think will be more favored by companies?
Fang Han: First of all, a "hundred-model war" situation has now formed. Everyone is making large-model bases. In the future, the large-model base segment will definitely be reduced to a few large manufacturers. More companies should be in a position to use large models for applications. Then I think there will be more and more talents developing applications based on large models.
Those who are engaged in the underlying training of large models, optimization of algorithms and architectures will gather in large companies or large model teams, but we believe that the biggest giants are not necessarily the large model companies themselves, but those companies that make strong applications based on large models. Once these companies grow up, they will also build their own large models.
We believe that "application is king", which means that companies that create new business models based on AI and big models will reap the biggest dividends. We believe that in the next ten years, there will be new giant companies like ByteDance, Meituan, and Didi, and they will definitely grow from 0 to 100. Companies founded this year or next year should have this possibility and opportunity.
